Every ranking we publish is denominated in bits, which is an unusual unit for a product data discussion and worth explaining rather than asserting.
The reason is that the question is genuinely about information. A shopper facing a shelf has a set of candidates and is trying to reduce it. Every fact they learn removes some of that uncertainty. Information theory has measured exactly that since 1948, and the unit it measures it in is the bit.
Why not a percentage complete
Completeness cannot compare unlike fields. Screen size at 60% and processor at 25% tells you which is emptier and nothing about which is worth filling. They are the same unit measuring different things.
Bits are comparable because the thing being measured is the same in both cases: how much does knowing this narrow the choice? An attribute recorded on every product with an identical value scores zero, correctly, because knowing it removes nothing. An attribute recorded on nothing scores zero today and can score very highly once populated, which is the distinction that matters for planning.
How many products a shopper is still choosing between, per bit learned
the definition of a bit, applied to a 216-product shelfOne bit halves the shelf. That fixed meaning is why screen size and processor can be compared at all.
View as table
| Value | |
|---|---|
| 0 bits | 216 left |
| 1 bit | 108 left |
| 2 bits | 54 left |
| 3 bits | 27 left |
The second half is demand
Information alone is not enough. A field could partition the catalogue perfectly and still be worthless if no purchase decision turns on it. So each attribute is also scored against the stated requirements for the category, and a field nobody's decision depends on scores near zero however neatly it divides the shelf.
This is where our ranking and a coverage ranking come apart. On the laptop shelf, processor was missing on 162 of 216 products, third by absence. Only 2 of 32 stated requirements mentioned a processor, which put it last by demand. Storage and processor swapped places, and that swap is the entire argument for measuring demand rather than absence.
Fixing the emptiest field first would have been the most expensive way to change nothing.
What we cannot stand behind
Screen size came first in 12 of 12 runs across every seed, weighting and imputation scope we tested. Positions two through four moved between runs. We publish the whole ranking because it is informative and we say plainly that only the top of it is stable.
There is also a real limitation in how we simulate a completed catalogue: we can only use values the shelf already records. Where an attribute's recorded values are unrepresentative of its true range, that simulation is wrong in a way we can name and cannot yet correct. It needs a value distribution from outside the catalogue being repaired.
The measurement and the fixture it ran against are both in our repository, and one command regenerates every figure above. That is the property we would most want from a vendor making claims about our data, so it is the one we built for.