Build or buy PIM for ecommerce: what a PIM is, when you need one and how to choose
Product data rarely breaks all at once. It drifts, one channel and one language at a time, until nobody is sure which spreadsheet is right. Here is how to tell when that drift needs its own system, and which kind.
Should you build or buy PIM for ecommerce?#
The right route follows from your catalogue: no PIM for one channel and one language, buy or self-host when channels multiply, and build only when the product data model is your edge.
Four routes are open, and each one suits a different catalogue. You can buy PIM software as a service, such as Salsify or Akeneo's cloud packages. Alternatively, open core or self-hosted software runs on your servers, such as Pimcore or Akeneo Community Edition. Then there is building your own. Finally, there is no PIM at all, just a spreadsheet plus your platform's metafields. Other vendors exist too, among them inriver, Pimberly and Plytix, listed alphabetically and not ranked.
Four facts about your catalogue decide between the routes:
- Channels and languages. One channel in one language points to no PIM. Each channel or language you add multiplies the values to keep correct, as the worked example below shows.
- Product shape. Standard families of attributes, such as size, colour and material, fit a vendor's model well. However, rules that define the product, such as which part fits which machine, may not.
- Hosting skills. A team that already runs PHP apps on its own servers can carry open core. In contrast, a team without one gets more from SaaS, where the vendor runs the servers.
- Where your edge is. If the way you describe products is what customers pay for, building can win. However, if the data model is a means, a vendor's model costs you little.
So measure the pressure first. Count your channels, your languages and the attributes on a typical product, then multiply them out. Next, compare that load with the limits the platforms publish, such as the 2,048 variants per product in Shopify's Help Center, accessed 5 October 2026. Then weigh the four routes on the same six axes, which a table further down does side by side.
You will meet vendor advice on this choice, and it is worth reading as what it is. For example, Akeneo, which sells a PIM, argues in its build or buy article that "for most organizations, there is a definitive answer", namely buying. However, that is a vendor's claim. The conditions above rest instead on published limits, licence terms and plain arithmetic.
What is a PIM, and what is it not?#
A PIM is the one system that holds, enriches and publishes every product attribute to each channel; it does not hold stock, prices or orders. Jorij Abraham's 2014 Springer book on the subject puts the trigger plainly. As companies sell online, they find that not all "information necessary to sell their products is available".
Because the gap is about information, not stock, a PIM has a narrow job. It stores the descriptive record: names, descriptions, attributes, sizes, materials, care notes, images and their translations. Then it shapes that record for each place you sell. The same book argues that "PIM is a new business process with its own unique implementation" and management challenges. As a result, buying software alone rarely fixes it.
It helps to name what a PIM is not:
Not an ERP,
because the ERP owns the SKU, cost, price, stock and orders. A PIM reads the SKU from it and adds the story.
Not a DAM,
because a digital asset manager stores and versions media files. A PIM links to those files and says which image belongs to which product.
Not the store catalogue,
because Shopify or another platform only shows products to shoppers. Its catalogue is one channel the PIM feeds, not the master copy.
The data model is the heart of any PIM. In Pimcore's documentation, version 2026.3, "Data Objects are the foundation of Pimcore" and its PIM features. There, a class definition sets which attributes each object has. Every PIM has some version of that idea. Also, how well that model fits your products decides most of the buy or build question.
What are the signs you have outgrown spreadsheets and your platform catalogue?#
Shopify allows 3 options and 2,048 variants per product (Shopify Help Center, accessed October 2026), so outgrowing it usually shows first as channels, languages and attributes. In practice, few stores hit those ceilings. Instead, they hit the point where the copies stop agreeing.
The Help Center puts the cap in one line. "You can create up to 2,048 variants for a product." On the same page, Shopify, accessed 5 October 2026, also says each product can have up to three options, such as size, colour and style.
| Limit | Per product |
|---|---|
| Options per product | 3 |
| Variants per product | 2,048 |
There is a store-wide throttle too, since Shopify's API limits page, accessed 5 October 2026, says that once a store has 500,000 variants, at most 10,000 new variants can be created per day. As a result, a large range loaded in one push can stall for days.
However, the signs that matter most are about drift, not size. Check your catalogue against this list:
- The same product reads differently on two channels, and the team cannot say which version is right.
- A new channel takes weeks of copy and paste, because each one wants its own fields and formats.
- Translations lag the source language, so one market sells last season's description.
- Attributes live in people's heads. Filters on the store miss products because a field was left blank.
- One spreadsheet has become several, each owned by a different team, with no record of who changed what.
If three or more of these are true, the pressure is real. Then the next question is what each channel demands.
What does each sales channel demand from one product record?#
Google Merchant Center caps a title at 150 characters and a description at 5,000 (Google, accessed October 2026), and every other channel sets its own rules. Because of that, one master record cannot be pasted everywhere as it stands.
Show data table
| Item | Value |
|---|---|
| ID | 50 |
| Title | 150 |
| Description | 5,000 |
| Additional image link (URL) | 2,000 |
A title gets 150 characters against 5,000 for a description, so one master record cannot be pasted into every feed as it stands.
For example, say a 200 character title works on your own store. It has to be cut to fit the 150 character limit in Google's product data specification, accessed 5 October 2026. Also, a marketplace may want the brand first, a size chart in its own format and its own category codes. Meanwhile your storefront wants long copy, swatches and filters.
This shaping is the real work a PIM automates for you. You keep one rich record, and then the PIM applies per channel rules: which fields to send, how to shorten them, which units and languages to use. Without it, each channel becomes its own spreadsheet. Then the copies drift apart as soon as someone edits one of them.
In short, a channel is a format with rules, not just a destination. So the number of channels, more than the number of products, is what turns copy and paste into a system problem.
Where does a PIM sit between ERP, storefront and marketplaces?#
The ERP owns the SKU, price and stock, the PIM owns descriptions, attributes and media, and the storefront and marketplaces read what each one publishes. Data flows one way for each field. Therefore that direction has to be agreed before anyone builds a connector.
In practice, the PIM vs ERP line is drawn in a field ownership map, a short table: each field, the system that owns it, and the systems that only read it. Price belongs to the ERP, so the PIM never edits it. A description belongs to the PIM, so the store never edits it. When two systems both write a field, the copies fight. Then the last write wins silently.
The vendors document how data gets in and out. Pimcore's Data Importer, version 2026.3, reads data from sources such as SQL, SFTP and HTTP. It takes formats including CSV, JSON and XLSX and writes them into data objects. Also, Salsify's integration guide, accessed 5 October 2026, names three ways out: a product CRUD API, a channel feed and webhooks. It notes that a product feed from a channel "is useful for mass exports or changes", while webhooks suit near real time updates. Finally, Akeneo's REST API overview, accessed 5 October 2026, uses JSON as its only data format.
Because each connector follows the field map, settle the map first. Then the connector list writes itself.
How many product field values would you have to keep correct?#
In an illustrative catalogue of 2,000 SKUs with 40 attributes, 3 channels and 2 languages, someone keeps 480,000 field values correct, and each added channel adds 160,000. The number is a model, not a survey, but it shows where the load comes from.
The sum is plain multiplication. In this illustrative case, 2,000 SKUs times 40 attributes is 80,000 values for one channel in one language. Then two languages double it to 160,000 per channel in the example. Finally, three channels, say your store, a shopping feed and one marketplace, make it 480,000 values in this example.
Your field values to keep correct
Enter your SKUs, attributes per product, sales channels and languages; it multiplies them and shows what one more channel adds.
Field values to keep correct
480,000
- Values one more channel adds
- 160,000
An illustrative model, modelled not measured.
Not every field changes per channel. Even so, every field has to be checked per channel, because one stale value on one marketplace is the one a shopper sees. Also, the model is linear in every input. Doubling languages costs as much as doubling channels.
Therefore channels are the input to watch. In this illustrative model, a store that adds a marketplace and a second language in one year has doubled its load twice. So run the numbers for your own catalogue first. For example, the routes compare differently at 80,000 values than at 800,000.
How do buying, open core, building and no PIM compare?#
Every route trades one thing for another: SaaS trades control for vendor upkeep, open core trades hosting work for source access, building trades upkeep for fit, and no PIM trades scale for simplicity.
The table sets all four routes on the same axes, with the case where each one wins and the cost it carries.
| Axis | Buy SaaS (Salsify, Akeneo cloud) | Open core or self-hosted (Pimcore, Akeneo Community) | Build your own | No PIM (spreadsheet plus metafields) |
|---|---|---|---|---|
| Cost drivers | A subscription set by quote | A licence for paid editions, plus your own hosting | Salaries for a standing team, plus hosting | Staff hours spent on each value, plus the store plan you already pay |
| Control | Within the vendor's data model, roadmap and API limits | Full source access, within the licence terms | Total: your model and your rules | Limited to the store's fields, such as 3 options per product on Shopify |
| Upkeep | The vendor runs hosting and upgrades | You run hosting, patches and upgrades | You own every feature, fix and upgrade | Manual checks on every channel |
| Time to value (reasoned) | Shortest to a working system, though the data work still takes its own time | Longer, because you install and host first | Longest, because features come before data | Immediate, since nothing is installed |
| Lock-in and exit | Vendor terms decide; data leaves through the API or channel feeds | Data sits in your own database, yet a vendor can end support for a self-hosted version | The least vendor lock-in, but the system depends on your own team's knowledge | Lowest, since a spreadsheet moves anywhere |
| Integration effort | Vendor connectors, within published rate limits | Your team builds on the vendor's APIs and importers | Every connector is built and kept by you | Uploads or apps per channel, repeated by hand |
| Wins when | Several channels, standard product families, no team to run servers | Several channels, a team that runs PHP apps, a need to hold data in-house | The product data model is your edge and a team will own it for years | One channel, one language, a modest range |
| Real downside | You work inside its model and limits, at a price set by quote | Hosting is yours, and commercial production use needs a paid edition | Every feature a vendor ships is yours to build and keep | Drift grows with every channel and language you add |
Source: Salsify pricing address, Akeneo compare packages and versions pages, Pimcore pricing and Community Edition pages, all accessed 5 October 2026. Time to value is reasoned from the steps each route needs, not a published figure.
Because the vendors publish different amounts, here are the same fields for each vendor named.
| Field | Akeneo | Salsify | Pimcore |
|---|---|---|---|
| How it runs | SaaS packages; Community Edition self-hosted | SaaS | Self-hosted editions, or PaaS |
| Licence | Commercial SaaS; Community Edition under the Open Software License 3.0 | Commercial SaaS | Pimcore Open Core License, every edition |
| Published price | None for its higher packages, which say "Call for quote"; Community Edition free | None; its pricing address opens the home page and a demo request | Professional $9,900 and Enterprise $29,900 a year; Community free under EUR/USD 5 million revenue |
| Published API limits | 4 concurrent calls per connection, 100 requests a second | 10,000 requests an hour per organisation | None published; on your own servers, your hosting sets the pace |
| Ways out | REST API, JSON only | Product API, channel feeds, webhooks | Datahub GraphQL, REST and webhooks |
Source: Akeneo compare packages page, Community Edition licence and API pages; Salsify pricing address and rate limiting page; Pimcore pricing, Community Edition and Datahub pages (version 2026.3); all accessed 5 October 2026.
The licence terms are where the routes differ most. Pimcore's pricing page, accessed 5 October 2026, lists Professional at $9,900 a year and Enterprise at $29,900 a year. Also, it says there are no limits on products, assets or users. Meanwhile Pimcore's Community Edition page, accessed the same day, makes that edition free for non-production use, non-profits and "companies under EUR/USD 5 million in revenue". By contrast, Akeneo's Community Edition repository carries the Open Software License 3.0. Neither Salsify nor Akeneo's higher packages print a price; Salsify's pricing address, checked on 5 October 2026, opens its home page instead.
Similarly, published rate limits differ by vendor. Salsify's rate limiting page, accessed 5 October 2026, sets 10,000 API requests per hour per organisation. Akeneo's limits appear in the implementation section below. However, Pimcore's Datahub documentation, version 2026.3, describes GraphQL, a simple REST API and webhooks, and publishes no request limit. On a self-hosted install, that means your own servers set the pace.
10,000 per hour
API requests per organisation
5,000 per hour
Session requests per user
Integration traffic has its own hourly budget, so read it as a design input for how fast you can sync.
| Option | requests per hour |
|---|---|
| API requests per organisation | 10,000 per hour |
| Session requests per user | 5,000 per hour |
Self-hosting has its own exit risk too. Akeneo's versions page, accessed 5 October 2026, gives 30 September 2026 as the end of support for version 7.0. After that date, it says Enterprise Edition users "will have to move to the SaaS version" to keep SLA-backed bug patches.
At its strongest, the build case runs like this. When your products are defined by rules rather than attributes, such as parts whose fit depends on a machine or a vehicle, the data model is the product. Because a vendor model stores attributes by product family, rules like these can fight it at every upgrade. Also, a built PIM runs at the pace your own servers allow, under terms only you can change, and stays supported as long as your team keeps it. Even Akeneo's own article grants that "one of the biggest benefits is control".
Building still costs the most upkeep, however. You own search, versioning, user roles, review workflows, imports, exports and every channel connector. Since each of those is a product in itself, a built PIM needs a standing team, not a one-off project. For the same trade-off across other software, see custom software vs off the shelf.
What does a PIM implementation actually involve?#
Implementation is mostly data work: agree the attribute model, clean and load the catalogue, then build connectors within limits such as Akeneo's 100 requests per second (Akeneo, accessed October 2026). The phases look much the same whether you buy or build.
- 1
Model.
Agree product families, attributes, required fields per channel and who owns each field. This is the step teams rush and later redo.
- 2
Clean.
Remove duplicates, fix units, fill blanks and settle one name for each value, which is where most of the time goes.
- 3
Load.
Import the cleaned catalogue in batches, check each batch and fix what failed before the next one.
- 4
Connect.
Build the ERP feed in and the channel feeds out, following the field ownership map.
- 5
Run.
Hand enrichment to the teams who own each field, with review steps and completeness rules per channel.
Connectors carry the most technical risk, because every system throttles. For instance, Akeneo's API good practices page, accessed 5 October 2026, allows 4 concurrent calls per connection and 10 per instance. It also allows 100 general requests per second and 3 attribute option writes per second.
| Limit | Value |
|---|---|
| Concurrent calls per PIM connection | 4 |
| Concurrent calls per PIM instance | 10 |
| General requests per second per instance | 100 |
| Attribute option writes per second per instance | 3 |
Akeneo says plainly that your integration "should anticipate this throttling" and handle failures. It also says to wait out a Retry-After header on a 429 response. Shopify's API limits page, accessed 5 October 2026, adds its own cap of 10,000 new variants per day once a store passes 500,000 variants. As a result, an initial load is planned in days and batches, never as one overnight push.
Finally, budget time for people, not only for systems. The 2014 Springer book's point holds, because a PIM is a business process first. The teams who write and check product data need roles, rules and the time to use them.
When is a PIM the wrong fit for your store?#
A PIM is the wrong fit when you sell on one channel in one language and a clean spreadsheet plus platform metafields already keeps every field correct. Run the earlier model with 1 channel and 1 language. Then the illustrative 2,000 SKU catalogue drops to 80,000 values.
At that size, one well kept sheet and one owner can hold the line. Also, Shopify's own catalogue does more than many stores use. Its variants page, accessed 5 October 2026, lets you connect category metafields to variant options. That way you can "reuse your data across products instead of having to recreate it each time".
Three cases call for something other than a PIM:
| Case | What to use instead |
|---|---|
| One channel, one language, a few hundred products. | Use the platform catalogue with metafields and a single source spreadsheet with clear column rules. |
| The pain is images, not attributes. | A DAM or a tidy shared media library may solve it on its own. |
| The pain is stock and price drift. | That is an ERP or inventory sync problem, and a PIM will not fix it. |
Revisit the decision when you add a second channel or a second language, since that is usually when the count, and the drift, start to climb.
Where to go next#
Treat the storefront, the integrations and the PIM as three separate decisions, and settle field ownership before choosing any of them. Each one rests on its own evidence.
If the storefront is the open question, headless vs traditional ecommerce compares the two shapes. For content that sits beside your products, types of CMS sets out the options. When product data meets stock and suppliers, supply chain management software covers that side.
For the store itself, see ecommerce development. For the feeds between your ERP, store and marketplaces, see API integrations. The vendor documentation linked above is enough on its own to run the test and draw your field map.