Mercur

Product Taxonomy for a Marketplace: Categories, Attributes, and Variants

Catalog and product data~15 min
Product Taxonomy for a Marketplace: Categories, Attributes, and Variants

A product taxonomy on a marketplace is the classification tree your commission rates, mandatory fields, and seller permissions all hang from, which is why changing it after launch is a migration.

A category tree looks like housekeeping. Give it a spreadsheet and a week, and it is done.

In fact, it is the skeleton that carries your commission rates, your field requirements, and your seller permissions. A skeleton is not something you edit after launch. You migrate it.

This article breaks down:

  • Why do classification and navigation need separate trees?
  • What does one extra level of a category tree cost?
  • Variant or separate product: which test settles it?
  • Why is changing the tree after launch a migration?

Key insights

  • Buyers barely touch your category tree. Three mechanisms read from it every day: which fields are mandatory, what commission applies, and who is allowed to sell there.
  • An optional field stays empty forever. A seller fills in the minimum validation, gets through, and moves on to the next thousand offers.
  • Moving 1,200 offers from an 8% node to a 12% node moves €12,000 a month from the sellers to you. That is a change of commercial terms with its own notice period.

Who is a marketplace category tree designed for?

Buyers almost never walk your classification tree. They have a menu, filters, a search box, and links from campaigns, and those usually look nothing like the tree, because they change every season.

You design the tree for three mechanisms that read from it:

  1. Field requirements. The node says which product information is mandatory and which list of values a seller may pick from.
  2. The commission rate. The grid hangs on the tree and inherits downward: you set a rate high up, then walk down and override selected branches.
  3. Rules and permissions. What may be sold, by whom, after what validation.

Hence the sentence to bring to your first workshop: every node is a promise that you can operate it. That means naming the fields that make sense inside it, assigning it a rate, and saying who sells there.

A node without those three answers is an empty label, and sellers will fill it with whatever they have.

And a less obvious consequence: all three mechanisms break when you change the tree later. This is not an aesthetic decision.

It is a decision about costs you will carry for years.

Why do classification and navigation need separate trees?

The cheapest thing you can do in this project is separate two ideas that blur into one whenever people talk about them.

Classification answers the question "what is this." One product has exactly one place. Field requirements, commission, and rules all read from it.

It changes rarely and expensively.

Navigation answers the question "how do I find this?" The same product can sit in many places at once: in a promotion, in a brand collection, in a filter result. It changes with every campaign.

Practitioners who build large catalogs say it plainly. The tree sellers list against does not have to be the same tree the storefront shows, and that separation is exactly what gives the operator control.

Merge the two, and you get a tree where "holiday gifts" stands next to "washing machines." The consequence is countable. A campaign node gets a commission rate and a set of mandatory fields, and once the season is over, nobody can say what you really earn on washing machines, because part of the turnover leaked into a node that was meant to live for six weeks.

Navigation is allowed to skip a product that does not fit the campaign. Classification has no such right: a product with no node has no fields, no rate, and no rules.

What does one extra level of a category tree cost?

The same distinction priced two ways: as a category it is 240 new nodes, 720 names to translate, 240 commission rates, 240 field decisions and about 80 hours, while as an attribute it is one field, three values, twelve strings and an hour and twenty

Assume a catalog with 900 leaves (the end categories) in 3 languages. Somebody proposes splitting 120 of those leaves into three subtypes each, because "it will be easier for the buyer." It looks like 120 clicks.

It gives you 360 leaves where 120 stood, which is 240 new nodes.

Each one has to be named in three languages: 720 text strings. Each one gets a commission rate: 240 assignments, in practice done by hand, because in some systems the rates in the grid cannot be set programmatically.

Each one needs its set of mandatory fields reviewed: 240 decisions. At twenty minutes per node, that comes to 80 hours.

That is two weeks of one person's work before the first offer goes live.

The same distinction made with an attribute is one field and three values: 12 strings across three languages, no rates at all, one decision about whether the field is required. Sixty times less work at launch and at every change after it.

Filtering comes free on top, because the storefront filters on attributes.

A working rule that settles most of these arguments: if the distinction exists so the buyer can narrow a choice, it is an attribute. If it drives an obligation or money, it is a category.

The control question is short. Does anything you sign with a seller change because of this distinction?

Depth comes with one more installment: every seller maps their own tree onto yours, and practitioners report that this is one of the most time-consuming parts of onboarding.

Which 3 classes of product field does a category node carry?

A product page in the "washing machines" category can carry 34 fields: 9 required to sell, 7 required for comparison, 18 optional. Make that split explicit, because each class has a different price.

Required to sell are the fields without which an offer has no right to appear. Their price is friction: every one of them is something the seller may not have in their file. Some of them are forced by regulation.

Required for comparison are the fields without which your filter is dead: capacity, energy class, dimensions. They do not block a sale, but without them the buyer cannot narrow anything down.

Optional fields deserve one honest sentence: an optional field stays empty forever. The seller fills in the minimum that validation lets through and moves on to the next thousand offers.

Practitioners describe this as the rule. A field that is optional "for now, we will ask later" is a field that does not exist.

Attribute inheritance down the tree cuts that work back: you define a set of fields once, high up, and the leaves inherit it and only add their own. This is how product information management systems model it.

Values from the parent level apply downward until somebody overrides them.

The trap is symmetrical to the benefit. An "Electronics" node can have 340 leaves and 210,000 offers under it.

Flipping one field on that node from optional to required means 210,000 offers to evaluate again. Here you need to know what your platform will do: nothing, flag them, or pull them off the storefront.

Catalog rules carry that question, but you ask it while you are designing the tree, because the answer decides how high up you are allowed to place a field.

Variant or separate product: which test settles it?

The test goes like this: does the buyer pick it in one move on the same page, without changing their mind about what they are looking for?

Size and color qualify. A different storage capacity usually qualifies.

A carton against a pallet qualifies too, although the shipping method changes with it. A different model from the same line does not qualify, because there the buyer is changing their mind about the product rather than about its version.

The product data standards that search engines consume settle it the same way: a variant group is a template that nobody buys, and the usual axes of variation are size, color, material, and pattern. This is a convention that external channels read.

Variants broken out into separate products blow the product page apart. A shoe in 6 sizes and 3 colors is 18 variants.

If 4 sellers list it, the correct model is one product page, 18 variants, and 72 offers. The wrong model is 72 product pages: 72 scattered sets of reviews and search positions, no way to tell the buyer "not in your size, but we have it in a 43," and the duplicate factory from duplicates and merging.

Separate products squeezed in as variants break the things a demo never shows. Availability is counted per variant, so a variant that is really a different product eats somebody else's stock.

Offer matching stops working. And the fields that regulation requires for one of those "variants" have nowhere to sit, because the page carries one set of fields.

Where should a new distinction go: category or attribute?

The same change has three different prices. It usually arrives as one sentence in a workshop: "let us tell three kinds of this product apart." Numbers as above: 120 leaves, 3 new values, 3 languages.

What does each option cost?

A new level in the tree

A new attribute

A new variant on the page

What it drives

fields, commission, permissions

filtering and comparison

a choice on one page

New objects to create (count)

240 nodes

1 field plus 3 values

3 variants per page

Names to maintain in 3 languages (count)

720

12

0 (uses an existing attribute)

Commission rates to assign (count)

240

0

0

Field sets to review (count)

240

1

0

What the buyer sees

a menu entry and its own page address

a filter and a column in the comparison

a selector on the page

Changing it after launch means

migrating offers, rates and addresses

revalidating offers

rebuilding pages

The recommendation: attribute by default, category only with a stated reason recorded on the node itself.

Why is changing a category tree after launch a migration?

The four things a recategorization moves at once: the commission rate on contracts already running, worth EUR 12,000 a month on the worked example, the required fields, what a seller is allowed to sell, and every page address

The point of this whole chapter: a recategorization moves four things at once.

1. The commission rate changes on contracts already running

You move 1,200 offers from an 8% node to a 12% node. On a €1,000 cart, the seller used to receive €920 and will now receive €880.

If those offers do €300,000 of turnover a month, you are moving €12,000 a month from their side of the table to yours. That is a change to commercial terms, and it has its own procedure and its own notice periods (the rules on decisions about sellers, and the documents you sign with a seller for the documents).

2. The required fields change

Some of the offers you moved become non-compliant at the moment they move.

3. What a seller may sell changes

Permissions hang on nodes, so a move can take a category away from somebody or, worse, open one they were never meant to enter (restricted and prohibited products).

4. Page addresses change

New paths, old ones to redirect, search rankings counted from zero again.

Hence the conclusion: a recategorization is a project with a plan, a window, and a rollback.

Industry product taxonomy or your own tree?

For many assortments, there are recognized public classification systems: the codes used in public procurement, cross-industry standards built on four levels and an eight-digit code that cover tens of thousands of classes and properties, and the fixed taxonomies that advertising channels publish.

Mapping to a standard buys you two things. Integrations with suppliers and comparison engines stop being a negotiation, because both sides speak one code.

And a seller who has mapped their catalog to that standard once does not do the work from scratch again.

And it costs you one thing: every mapping is lossy. The standard does not know the distinction you make money on, or it knows three where you need one.

External channels assign a category automatically when you do not supply one. So "we are not mapping" does not mean "we are not mapped." It means "somebody will do it for us."

The recommendation: your own tree as the operational classification, the standard code as a field on the node. Framing this as either/or is false and expensive.

What does the category tree change about the rest of your catalog?

1. Catalog rules enforce the structure you build here

Product approval runs through that structure (the approval queue).

2. The number of leaves is a fixed cost

It comes back with every new language (a multilingual catalog), with every change to the commission grid, and with every new seller who maps their tree onto yours.

3. Field requirements are a policy

Field ownership covers who has the right to fill a field in and to change it. Catalog debt covers what to do with what is already empty.

4. A tree nobody cleans turns into debt

Nodes with no offers and no owner do not disappear on their own, so keep the name of a responsible person against each one.

How do you check one category node in 90 minutes?

Do not audit the tree. Take one leaf with at least 200 offers and 5 sellers.

  • Count the fields and split them into the three classes: required to sell, required for comparison, optional.
  • Take a sample of 50 random offers and count how many have even one optional field filled in. That percentage closes the discussion about whether "we will ask the sellers later."
  • Count how many product pages are really variants of the same thing: the same product in a different size, listed as a page of its own.
  • Check how many offers carry a rate other than the node rate (exceptions negotiated per seller). That tells you what moving the leaf costs.
  • On a test environment, move the leaf under a different node and write down what changed by itself: the rate, the fields, the permissions, the address. Everything else will be manual work.

Three questions for your platform vendor, with a request to show it on screen:

  1. You flip a field from optional to required on a node with several hundred leaves. What happens to the offers already published: nothing, a warning, or removal from the storefront?
  2. You move a thousand offers into a different category. Show what changes automatically and what has to be finished by hand.
  3. The classification tree is independent of the menu. Show whether navigation can be rebuilt without touching the classification.

Which mistakes do operators make about product taxonomy?

1) A tree designed as a menu

Commission and mandatory fields hang on a node that was meant to live for six weeks, and after the season nobody can work out the profitability of the real category.

2) Depth instead of an attribute

Several hundred names to translate, several hundred rates, and just as many field sets.

The filter on the storefront still has to be built separately.

3) An optional field as the compromise reached in the workshop

The field never gets filled in; the filter and the comparison are dead, and a year later you go back to those offers asking sellers to complete them.

4) Variants listed as separate products

Scattered reviews, no information about other sizes, offers that never merge onto one page, and a steady production line of duplicates.

5) A recategorization treated as an edit

A silent change to commercial terms, part of the catalog suddenly non-compliant and page addresses broken.

All of it at once, with no window and no rollback.

What do you still have to settle about your own tree?

This is a method for designing a structure. It is not a finished tree for your industry.

You cannot copy the shape of your categories from anybody. It follows from where your rates, your obligations, and your permissions change, and that knowledge sits with your commercial team.

What is deliberately absent here: enforcing the structure with rules, the quality of descriptions (what you can enforce on content), translating names and values, field ownership, matching on attributes, and cleaning up what is already broken. The commission grid and the mechanics of search have chapters of their own.

There is also no credible number for the "optimal depth of a tree," and we are not going to pretend that there is. You calculate it on your own data: how many commission rates you have, how many sets of mandatory fields, and how many levels of permissions.

That is how many branchings you need. Everything else is an attribute.

Confirm with a lawyer which fields regulations require in your categories. We say only this much: the higher in the tree you place such a field, the more offers its introduction will touch.

Summary: What does a marketplace category tree have to carry?

Three things per node, or the node is an empty label: a set of fields that makes sense there, a commission rate, and a list of who may sell into it. Everything else the buyer needs belongs to navigation, which is allowed to change every season because nothing hangs off it.

The working rule for a new distinction is short. If it exists so the buyer can narrow a choice, it is an attribute.

If it drives an obligation or money, it is a category. Get that wrong in the cheap direction, and you pay 80 hours for what a single field would have done.

Take one leaf with 200 offers, sample 50 of them, and count how many have a single optional field filled in. Building a marketplace where a rate, a field set, and a permission all hang off the same node?

Talk to us about the build.

Frequently asked questions on marketplace product taxonomy

What is a product taxonomy on a marketplace?

The classification tree that says what a product is, and the thing your mandatory fields, commission rates, and seller permissions all read from. It is separate from navigation, which answers how a buyer finds something and may put the same product in a promotion, a brand collection, and a filter result at once.

Should a distinction be a category or an attribute?

An attribute, unless it changes an obligation or the money. The control question takes a second: does anything you sign with a seller change because of this distinction?

A new category costs names in every language, a rate, and a field review. An attribute costs one field and its values.

When is something a variant rather than a separate product?

When the buyer picks it in one move on the same page without changing their mind about what they are looking for. Size and color qualify; storage capacity usually does; a different model from the same line does not.

Breaking variants into separate products scatters reviews and manufactures duplicates.

Ready to build?

If you have a tree design in front of you and want to walk through which distinctions have to be categories and which are fine as attributes, let's talk.