Product Data Quality on a Marketplace: What You Can Enforce

Product data quality on a marketplace is whatever a rule on a field can check and refuse, which is why it is settled by what you make machine-checkable rather than by what you ask sellers for.
Your own catalog was built over years by a team with a process and a standard of its own. A seller catalog is different: everyone does exactly as much as the publication gate lets through.
Description quality is therefore your design decision.
This article breaks down:
- Which product data rules can a machine check?
- Why render a product description from attributes?
- How do you measure completeness per seller?
- Which product fields may automation fill in?
Key insights
- "Better descriptions" is a programme with no finish line. "Eleven attributes and three photos in this category" is a programme with a date on it.
- Stop asking sellers to write. If eleven fields are mandatory anyway, render the description from them. The seller does less work, and every offer reads the same way.
- Photos are the cheapest win: minimum size, product filling most of the frame, plain background, no watermark. A machine checks every one of those, and the seller only uploads a file they already have.
- Show each seller their completeness score with the list of gaps. A seller will fill a field in for their own visibility long before they do it because you asked.
What can a marketplace enforce about product content?
A quality rule in a catalog is a function: it looks at the value in a field and returns "passes" or "does not pass". Nothing more.
A requirement you cannot write as a condition on a field is not a requirement. It is a request, and in a catalog with two hundred thousand offers nobody enforces requests.
The line between them runs exactly where the machine stops being able to check. That is why the fight over description quality is always lost at the same point: the description field is an empty box that accepts any text.
In it you can require presence, length, language, and the absence of code. You cannot require that it be a good description, because "good" is not a property of a piece of text.
It is a judgment about whether that text helped somebody. Practitioners on large rollouts call this the worst pain of the layer: there is no technical way to impose a description template on a seller, and terms carrying that template are a document rather than a validator.
Even accessibility standards, which have spent years trying to define "good alternative text for an image", end in the same place: whether the text conveys the sense of the picture is a judgment about context.
The consequence is uncomfortable, and better said to the board early than late. A quality program whose goal is "better descriptions" has no condition for being finished.
A program whose goal is "every offer in this category has eleven attributes and three photos that meet the standard" ends on a date, and you can be held to it.
Which product data rules can a machine check?

You will enforce whatever a machine can count or match against a pattern. The presence of a field and its format.
Agreement between the value, the category dictionary, and the unit. Text length ranges.
The shopping channels you feed anyway state these as plain numbers: a title from one to a hundred and fifty characters, a description up to a few thousand. The absence of banned patterns: HTML code, an email address or a phone number, a link to the seller's own store, a promise about the delivery date, all caps, decorative characters, a foreign alphabet.
The language of the field. Uniqueness of the identifier.
And completeness: the share of required fields in a category that are filled in.
You will enforce nothing that is a judgment about content. Whether the description is useful.
Whether it is true. Whether it answers the question the buyer arrived with.
Whether the tone fits your brand. No rule checks any of that, because it takes knowledge about the product that the field does not contain.
Say this without regret. The second list is not a problem to be solved.
It is a limit to be accepted. The difference between an operator who has a catalog after two years and one who has a garbage heap is that the first one stopped building the program around the second list.
Why render a product description from attributes?

If you already validate attributes, you already have everything a good technical description is made of. You just keep it somewhere else on the same product page.
Turn the direction around: do not ask the seller for text; render it out of the fields they must fill in anyway.
What this looks like in practice. Your washing machine category has twenty-four attributes, eleven of them mandatory.
The category template assembles those eleven fields into four paragraphs and a table of parameters, in a fixed order, with units written the same way in every offer. The seller fills in eleven fields they could never have skipped, and writes not a single sentence.
That changes the nature of the problem. The description stops being content somebody writes and becomes a view of data you already have.
The uniformity no set of terms will ever buy you arrives as a side effect. You are not controlling the text, only the template.
And the seller has less work: this is the only version of a quality requirement we know of that does not reduce the inflow of offers.
The scale is bearable as long as you do not start with everything. With nine hundred categories, you do not need nine hundred templates.
Twelve categories usually carry about four-fifths of the offers, so the first round is twelve templates.
Where this will not work: fashion, handicraft, books: anything bought for the impression it makes.
There, leave a field for the seller's own prose, as an addition under the rendered part.
Which product photo rules can a marketplace enforce?

Photos are the exception here, because almost the entire requirement is machine-checkable. The market standard is published in numbers: a minimum size on the order of five hundred by five hundred pixels, the product filling seventy-five to ninety percent of the frame, a plain or transparent background, no frames, no watermark, no price, and no call to buy burned into the image, no placeholder shots.
Add the number of photos per offer and the file formats. Every one of these is checked without a human.
The cost on the seller's side is as low as it gets, because you are not asking them to write, only to upload a file they usually have already. This is the best ratio of effect to cost in the whole quality layer, and the first thing worth enforcing.
How much of this you have. Assume two hundred thousand offers, of which thirty-eight thousand carry exactly one photo.
That is nineteen percent. The number sounds like a migration until you ask about the distribution: if they belong to a hundred and forty sellers, you are facing a hundred and forty conversations.
The distribution matters more here than the total.
Usability research shows two things at once: users look closely at photos that carry information about the product and skip decorative ones, but on listing pages most of the looking time goes to text. So treat photos as the cheapest requirement to introduce, and keep asking for the data.
And one symmetry that is easy to forget: you are checking the frame. No rule will tell you whether the photo shows that product.
How do you measure product data completeness per seller?

Completeness is the only quality measure you can compute honestly, because it measures filling rather than sense. Define it as the share of the fields required and recommended in a category, counted per offer, per seller, and separately per language.
An example. A seller has a hundred offers: sixty of them filled in at a hundred percent, forty at fifty-five.
Their score is (60 × 100% + 40 × 55%) ÷ 100 = 82%. That number, shown in their own panel with the list of missing fields and their standing against other sellers in the category, does more than three emails asking for better data.
The real leverage starts somewhere else, when completeness stops being information and becomes a parameter of visibility. For example: ninety percent and up with no restrictions, seventy and up with no right to win the product page, below that threshold only on the seller profile.
It works because it works on self-interest.
There is one price here, and you have to pay it deliberately. If completeness affects who wins the product page or a position on a list, you are obliged to disclose that to your sellers.
Ranking parameters and justifying decisions are led by "P2B Regulation on a Marketplace: Ranking, Terms Changes, and Suspension", and the rules that pick the winning offer belong to the chapter on the BuyBox.
And a warning from practice: the seller will work out faster than you that a field filled in with a single dash is formally complete. A score with no validation of the value measures skill at getting around the score.
Which product fields may automation fill in?
Where automation genuinely helps. Normalization: letter case, units, stripping decorative characters, and code.
Pulling attributes out of prose the seller sent you anyway. Out of fifty thousand offers with no "power" attribute, the value can be read from the description in about thirty-one thousand, just under two-thirds, and the rest you leave empty.
Translation of technical text. Spotting a description that does not match its category: the name says "rain boots", the tree says "umbrellas".
No rule on a single field catches that.
Where it is dangerous. In fields whose content you answer for in front of an inspector or the buyer.
The mechanism is well described in the literature on text generation: fluency and faithfulness to the input data are two independent properties of a system, so a model will fill an empty field with a smooth sentence it had nothing to build from. It does not know a fact the seller never supplied.
Hence a rule worth one sentence in your catalog policy: automation on form only. In practice, that means two queues.
Formatting changes go through on their own. Changes that add content go as a proposal for the data owner to confirm, and who that owner is gets settled by "Who Owns Product Data on a Marketplace?".
Fields about product safety and fields required by law are never filled in by a machine; their scope is covered by product safety and recalls and restricted and prohibited products.
Which 4 levers enforce product data quality?

The order of the columns is also the order of rising effectiveness, and of rising cost to implement.
What does product data quality change about the rest of your catalog?
1. Your attribute structure stops being a technical topic
If the description is a view of data, the category tree and the value dictionaries are a precondition for content quality. That is the category tree.
2. Who owns the value in a field matters more
A rendered description is only as good as its data, so an argument about data stops being an argument about taste.
3. Several languages get cheaper or dearer
What you translate is the template and the dictionaries. That is a multilingual catalog.
4. A bad description comes back as money
A buyer who was misled returns the goods: a €1,000 cart at a 12% commission takes €120 of revenue with it and opens the question of fault. Who pays for a refund and the commission on a refunded order takes that up.
5. Where the neighbouring topics live
Five neighbouring decisions have chapters of their own: the mechanics of rules, the approval queue, a backlog with no owner, duplicated product pages, and discount messages.
In which order do you introduce product data quality rules?
The order matters more here than the set of requirements, because a requirement introduced too early closes the door on sellers you do not have yet.
- Photos and mandatory attributes. Hard, measurable, cheap for the seller. Start with six weeks in warning mode: the offer stays, it gets a flag and a list of gaps. Only then do you reject at the gate.
- Description templates in the categories that carry most offers. Twelve of them. This is where you stop asking for text.
- A completeness score, with no consequences at first. Per offer and per seller, with the list of gaps and a comparison against the category. Two months to get used to the number.
- Only now, sanctions. First, the effect on visibility, disclosed in advance, then rejection. A sanction as step one costs you supply and buys you no quality.
Three questions for the platform vendor, with a request to show it on screen rather than in a feature matrix:
- Can the description be rendered from attributes per category, or is it only a free text field?
- Does the seller see the completeness score in their own panel, and can it be wired into the visibility of the offer?
- Can a quality rule reach offers that are already published, and does it flag them or take them off the storefront?
Which mistakes do operators make about product data quality?
1. A description template written into your terms instead of into a field
A document is not a validator: a year later you have a clause nobody cites and a catalog that has not changed.
2. A quality requirement introduced as extra work for the seller
Every field filled in by hand with nothing in it for them lowers the inflow of offers.
3. Completeness measured without validating the values
The score goes up; the quality does not. A dash typed into a field is also filled in.
4. Completeness affecting the ranking without disclosure
You are changing seller revenue with an undisclosed parameter; that is an argument you do not want to have after the fact.
5. Automation let loose on fields that state facts
You get smooth sentences with no cover in the data, and you answer for them with your own company, because it is your site that shows them.
What do you still have to settle about content quality?
It does not settle how much conversion a complete description adds. There is no credible market number here.
It depends on the category, the price, and what the buyer is comparing. Measuring it on your own numbers is cheaper than it looks: take one category, render descriptions from attributes for half the offers, leave the other half untouched, and compare after a quarter.
That is your number, and it is the only one that means anything in front of a board.
Two things lie outside our competence. Which fields on a product page are mandatory for you by law?
That depends on the category and the market. And whether automatically generated content has to be labeled, plus who answers to the buyer for whether it is true.
That is a question for a lawyer and for your contract with the seller.
Our own claim is narrower and does not depend on either answer: description quality is won by changing what a description is.
Summary: What decides product data quality on a marketplace?
What you turned into a condition on a field, and nothing else. Two lists settle the whole subject: the things a machine can count or match, and the things that take knowledge the field does not contain.
Build the programme on the first list, and it finishes on a date. The strongest move in it is to stop asking for text at all and render the description from attributes the seller has to supply anyway, which is the one quality requirement that leaves them with less work than before.
Then let the completeness score do the arguing, in their own panel, against the others in their category.
Count how many of your offers carry exactly one photo, then count how many sellers those belong to. Building a marketplace where a category template turns eleven fields into a description nobody has to write?
Frequently asked questions on marketplace product data quality
What product data quality can a marketplace enforce?
Presence, format, length, language, units, dictionary values, banned patterns, and completeness. Whether a description is useful, true, or on brand takes knowledge the field does not hold, so no rule reaches it.
That limit is worth stating to a board early rather than late.
Should sellers write product descriptions on a marketplace?
In most categories, no. Render the description from the attributes they already have to fill in.
Twelve categories usually carry about four-fifths of the offers, so the first round is twelve templates. Leave free prose where the purchase is driven by impression: fashion, handicraft, books.
How do you measure product data completeness?
As the share of a category's required and recommended fields that are filled in, counted per offer, per seller, and per language. Validate the values as well as their presence, because a seller will work out faster than you that a single dash counts as filled in.
Ready to build?
If you want to check how many of your quality requirements can be written as a condition on a field at all, let's talk.