Mercur

Faceted Search on a Marketplace: Filtering Across Sellers' Offers

Storefront and buyer experience~15 min
Faceted Search on a Marketplace: Filtering Across Sellers' Offers

Faceted search on a marketplace is the filter panel beside a listing, where each facet is a product attribute the buyer can narrow by.

Search is the one place where a buyer tells you outright what they want. On a marketplace, the quality of the answer does not depend on the engine you buy, but on how many fields your sellers filled in.

This article breaks down:

  • Should faceted search index the product or the offer?
  • What does an empty attribute do to a facet?
  • Which number do you sort by price on?
  • Why is a zero-results screen so expensive?

Key insights

  • Faceted search should index products where several sellers list the same thing, and offers where every item is one of a kind.
  • An empty attribute hides the offer from that facet entirely, so either require the attribute in the category or accept that part of your supply is unreachable.
  • Sort by price on the amount the buyer will pay, delivery included, so the cheapest result is the one that costs least.
  • A zero-results screen is expensive because it is a buyer naming what you do not stock, so log the query and use it to decide which sellers to recruit.

What does faceted search really search through?

A search engine does not read your catalog. It reads an index: a copy of selected fields, arranged so that you can search and filter on them.

In your own store, that is a technical detail, because your team fills the fields in. On a marketplace, it changes who owns the problem.

How complete the index is comes down to the decisions of a few dozen outside companies, each of which filled in as much as the form forced them to.

Practitioners see the same thing in every implementation: an offer from a seller has the required attributes filled in and almost none of the optional ones. Not out of ill will.

It is arithmetic. A seller posts the same assortment across several channels and fills in what those channels have in common.

Engines do not guess. An offer with no value in a field does not, by default, appear in the results of a filter on that field, and it is not counted in the counter next to it.

That is standard behavior for this class of tool. Two independent, widely used engines document separate parameters whose only job is to count the missing values or substitute a placeholder label.

Those parameters exist precisely because, by default, an offer like that is not in the result.

A filter that forty percent of your offers never filled in does not narrow the choice: it hides the assortment you paid for with seller recruitment.

Should faceted search index the product or the offer?

Product index returns 400 results across 20 pages; offer index returns 1,200 across 60 pages for the same supply.

It usually gets made in the first week of work, and almost nobody puts it as a question. Is a result on the list a product (with one representative offer), or an offer, which means the same product five times over?

Take 20,000 products and an average of three offers each, so 60,000 offers. The query "hair dryer" matches 400 products.

Index products and you get 400 results, or 20 pages of 20 items. Index offers and you get 1,200 results across 60 pages.

The same supply, a list three times longer. The first page looks worse still: if one product has five offers, twenty slots can show four different products.

The consequence for your metrics is the sneaky part. "Number of results" and "click-through on a result" mean different things in the two arrangements, so a comparison with the period before you changed the index stops making sense, and nobody announces that.

There is an economic argument as well. In practice, the real choice comes down to two or three offers.

You do not need a list that can set twenty offers side by side. You need a list that does not multiply the same product.

The default answer is a product index carrying the best available terms. An offer index makes sense where the offers differ on a feature people search by: used condition, model year, language version.

What does an empty attribute do to a facet?

Comparison of hard filter, soft filter and an explicit 'not specified' value, with hidden and non-matching offer counts.

Take a category with 12,000 offers and a "width" filter. The attribute is filled in on 60% of the offers, which is 7,200; 4,800 offers have no value there at all.

The buyer clicks "60 cm."

Among the offers with that field filled in, 2,100 match, and that is the list they see. But those 4,800 offers with no value hold roughly 1,400 matches of their own, at the same distribution as the filled-in ones.

A filter meant to narrow the choice to 2,100 offers has hidden 1,400 that meet the criterion: two-fifths of the genuinely matching supply, invisible to a buyer who has just said they want to buy it.

The opposite answer is no better. If offers with no value stay in the results, the list runs to 6,900 items, of which 3,400 do not meet the criterion.

Half the list is noise, and after the second disappointment, the buyer stops filtering. So this is a decision per attribute, taken deliberately.

What does each answer cost?

Hard filter

Soft filter

An explicit "not specified" value

Offer with no value

drops out of the results

stays in the results

goes into a filter option of its own

Results for "width 60 cm" (offers)

2,100

6,900

2,100, plus 4,800 on request

Matching offers that stay hidden (offers)

about 1,400

0

0, if the buyer expands it

Non-matching offers in the results (offers)

0

about 3,400

0

Who pays for the missing data

your supply

the buyer's attention

the buyer, with one click

What it fits

features that decide the purchase

supporting features

attributes with low coverage

A practical rule. Filter hard on anything where a mistake ends in a return: size, width, compatibility, power rating, class.

Filter soft on anything that is a preference. And if coverage of an attribute sits below roughly two-thirds, do not turn the filter on at all: it advertises a choice you cannot serve.

New requirements can rarely be forced onto a catalog that is already published, so coverage grows more slowly than the plan assumes.

Sorting by price: which number do you sort on?

Three price-sort definitions place the same product at €850, €890 or €870.

Three offers on one product. A: €890 plus €19 shipping.

B: €850 plus €39. C: €870 with shipping included.

A "cheapest first" sort places this product at €850, €890, or €870. Which one depends on the definition you pick, and that moves its position against every other product.

Three definitions, three different objections. From the product's lowest offer: the list looks attractive, but the buyer clicks on €850 and pays €889 once shipping is added.

From the default offer: the number is one you can buy at, except the product looks more expensive than it is, and seller B asks why they bothered cutting the price. From the total price including shipping: closest to the truth, and C wins.

Then the next question arrives at once. What shipping do you assume, when its cost depends on the address and on the rest of the cart?

Practitioners lean toward the total price: the assumption only has to be named and shown to the buyer (showing the price).

This is a market fact you need before you sign a contract, because it is a budget line and not a setting. In some implementations, search and sorting by price do not cover seller offers at all, because the index comes from the store layer and knows one price per product.

Its own. A product priced at €1,100 in the store, with three offers starting at €850, then sits in the sort at €1,100: the buyer never reaches your cheapest supply.

The cure is an offer index of your own. With 40 sellers at 1,500 offers each, that is 60,000 documents, and if every one of them changes prices twice a day, you get 120,000 updates a day, which is more than one every second.

There is also a question of authority: who adds a new facet, and how fast. We know of systems where the operator cannot do it themselves.

They file a request with the vendor, and repeat requests get turned down because rebuilding the index costs too much. At three new categories a year, that is three quarters without filters in exactly the places where you are opening up supply.

Why is a zero-results screen the most expensive on your site?

Four sources of zero-result queries, and the recovery screen that names the filter it dropped.

At 200,000 queries a month and 3% of them returning nothing, you have 6,000 sessions in which a buyer said outright what they wanted and got nothing back. That is the cheapest demand research there is, and it usually goes in the bin.

There are four sources.

  1. Inflection and synonyms: the core of a multilingual catalog.
  2. Typos.
  3. Filters that exclude one another: "brand X" has 800 offers, but not one of them has a color filled in, so adding "graphite" returns zero.
  4. And a hard filter on an attribute with low coverage, which is the mechanism from the previous section.

What to show instead of a blank: the same results with the most recently added filter removed, along with its name and the number of offers that come back once it is gone ("without the 'color: graphite' filter we have 800 offers"). That single change saves the session, because it does not make the buyer guess what they did wrong.

Independent usability research on e-commerce search points the same way. Queries that describe a product feature perform worst, and those are exactly the ones that run into unfilled fields.

Separately: boosting selected offers carries the same risk as a thumb on the scale in the rule that picks the winning offer. It is often unintentional, and regulations on transparency in platform-to-seller relations give a seller the right to know the main factors behind the ordering (the rules on decisions about sellers).

One number is worth measuring here: how many of the twenty slots on the first page are there because of a manual rule. Six out of twenty is 30% of your exposure handed out outside the algorithm.

Which 3 numbers does search give you about catalog quality?

Three monthly search metrics, and the arithmetic of invisible supply: 14% of 60,000 offers equals 8,400 items.

1. The share of queries with no results

Watch the trend. It goes up whenever supply grows faster than the data about it.

2. The share of clicks on the first position

Research on click data shows that position alone has a strong effect on how many clicks a result gets, regardless of how relevant it is. If 55 clicks out of 100 land on the first one, then your ranking is choosing for the buyer.

And the choice is made by a rule nobody in the company owns.

3. The share of offers no facet can reach

For every offer, check whether you can get to it with a query on the name or the identifier, through a category path, or through any combination of filters. Offers that fail all three routes are invisible supply.

At 60,000 offers, 14% is 8,400 items. That is as much as six sellers bring in at 1,400 offers each.

A €1,000 cart at a 12% commission is €120 of your revenue and €880 for the seller, and an offer with no route to it generates neither of those two numbers.

This is the warning of the article, and it runs against intuition: search quality gets worse exactly when the business grows. Every new seller brings in offers with the minimum set of data, so attribute coverage falls with every additional thousand offers, and the share of supply a filter can reach falls with it.

The curve goes down while the revenue chart goes up. Nobody connects the two, because a different team reports each one.

What does faceted search change about the rest of your catalog?

Which attributes are mandatory stops being a catalog question and become a storefront question. The list of required fields per category follows from which filters you want to be hard.

That changes the order of the conversation with the category tree and what you can enforce on content, and it is the only way coverage ever grows.

1. The index is a separate system with a separate owner

It has its own lag, its own outages, and its own budget. A price that shows in the seller panel after a second and in the results after ten minutes is normal behavior, but it has to be a decision and not a discovery.

2. Three neighbors inherit from this decision

The choice of the winning offer within a product (the rule that picks the winning offer), the comparability of the totals a buyer pays, and the visibility of pages with many offers in external search engines.

How do you check faceted search on your own catalog?

On a feature checklist, search is a single line item, which is why the gap only shows up after go-live.

  1. Run one query against a product with five offers. Does it return one row or five?
  2. Sort ascending by price on a product whose cheapest offer belongs to a seller. If the product does not jump to the position that offer implies, the index does not know about seller offers, and you have one of your own to build.
  3. Turn on a filter for an attribute that 40% of the offers do not have. Show the counter next to it and say where those offers went.
  4. Who adds a new facet, in what time, and how many times a year? "You file a request with us" is an answer. You just need to know its price and its limit.
  5. Show the zero results screen and the report of queries with no results from the past month.
  6. Show the list of offers that no query and no filter lead to. If there is no such report, ask for a query against the database.
  7. Does the information that a position has been boosted come back in the system's answer to a request for the list? Or is it drawn in by the template? The version for paid placement is covered by price history and discount messages.

1. A filter turned on for an attribute the catalog does not hold

Consequence: two-fifths of the matching supply is invisible at exactly the moment the buyer was closest to a decision.

2. A search index copied from the store layer

Consequence: the cheapest supply is unreachable through sorting, and sellers ask why they cut their prices at all. Then they stop cutting them.

3. Facets designed once, at launch

Consequence: the category with the largest new assortment starts with no filters, sometimes for quarters at a time.

4. Searches with zero results and no log

Consequence: you make seller recruitment decisions from conviction rather than from demand.

5. Sorting by price without the cost of delivery

Consequence: the buyer pays more than they would on a result standing lower down. And you lose the seller who chose not to charge for delivery.

What do you still have to settle about your own index?

This chapter does not settle relevance weights or the tuning of results. That is work for search specialists on your own data. It also does not settle inflection, synonyms, and per-language indexes, the shape of the tree and the value dictionaries, completeness of the description as a requirement on the seller, the choice of the winning offer on the product page, price presentation, or visibility in external search engines.

The most important number in this subject is one you have to measure yourself: what your coverage is on every attribute you want to filter by. That is a single query against the database, doable in one day, and the cheapest thing you can do before you switch the filters on.

Two obligations appear here as a product consequence: disclosing the main parameters that determine ordering, and labeling positions that stand higher because somebody paid for them. The families of regulation to name for your lawyer, without settling which of them apply to you:

  • those on transparency in platform-to-seller relations, P2B in market shorthand, in the part about disclosing ranking parameters
  • those on digital services, DSA in market shorthand, in the part about recommender systems and the labeling of advertising

The detail toward the seller is carried by the rules on decisions about sellers. Toward the buyer, price history and discount messages.

Our own claim is narrower and independent of those answers: a platform that cannot produce the list of offers no filter can reach does not know how much supply it has bought.

Summary: What does faceted search rest on?

Attributes, and nothing else. A facet is a field with values a buyer can narrow by, so every gap in the catalog becomes a gap in the panel, and an offer with an empty attribute is supply nobody can reach.

Two decisions sit above that. Whether the index holds products or offers, because it changes what a result even means.

And which number sorting by price uses, because the goods alone and the goods with delivery put different sellers at the top. Then let the thing report back: queries with no results, clicks on the first position, and offers no facet can reach are three numbers about your catalog that cost nothing to collect.

Run your ten most common filters and count how many offers they can never return. Building a marketplace where a facet is a field with a dictionary rather than a guess at somebody's prose?

Talk to us about the build.

What is faceted search on a marketplace?

Faceted search is the filter panel beside a listing, where each facet is a product attribute and its values, usually with a count of matching items. Facets can be combined, and each one narrows the set, which is why they only work on attributes held as fields with dictionaries rather than as free text.

Should faceted search index products or offers?

Faceted search should index products where several sellers list the same thing, and offers where they do not. Indexing products gives one result per thing with a best price attached; indexing offers gives one row per seller and a result count that surprises the buyer. The choice follows the catalog model.

Why does a filter return fewer offers than the catalog holds?

A filter returns fewer offers than the catalog holds because an offer with that attribute empty cannot match any value of it. The offer is live and unreachable, and nobody is told: the buyer sees a shorter list, the seller sees no sales, and the report shows the catalog as complete.

Ready to build?

If you want to count how many of your offers no query and no filter ever returns, let's talk.