What we cannot claim yet
We are early. These are the limits of the measurement and the gaps in our evidence, written down so you can weigh them before a pilot rather than discover them during one.
Four questions a buyer should ask us
2/4of the four are a no. They are on this page rather than in a procurement call because you would find out either way.
Where you are
12/12
Runs agreeing on the top repair
across every seed and weighting we tested
0
Live tests on merchant traffic
so there is no conversion figure, here or anywhere
0
Customers we can name
when that changes we will say so, with permission
The first number is why the measurement is worth buying. The second and third are why this page exists.
You will find these things eventually. A technical reviewer will ask how stable the ranking is, procurement will ask for the customer reference, and somebody will ask what happened when the method did not work. Those answers exist whether or not we publish them.
Discovering them at that stage is expensive for both of us: it stalls a deal that was going well and it makes everything else we said look like it was chosen carefully. Publishing them costs us some early conversations with people who need a safer vendor, which is the correct outcome for both parties.
So this is the whole list, taken from our own research log rather than rewritten for the website. It includes the run that failed and the reason it failed.
Anyone quoting you a conversion number before running a test on your traffic is quoting someone else's store.
What we will not do
Every vendor in this category leads with an uplift figure. Here is the arithmetic behind ours being absent.
The shelf we measured, before the repair
216
What it looks like
Twelve runs across every seed and weighting we tested. Only the first position held.
Only the first position is stable. We publish the rest and say so.
What happens
Screen size came first in 12 of 12 runs across every seed, weighting and imputation scope we tested. The top of the ranking is the part we stand behind.
They moved between runs. We publish the full ranking because it is informative, and we say plainly that only the first position is stable.
We have run no live experiment on merchant traffic. There is no number, and we will not produce one from a model.
What you get
No live test has been run on merchant traffic, so we have no conversion or revenue result. What we will do is show you the experiment design, with a held-back control group and a stopping rule fixed in advance, and report the outcome including when it is inconclusive.
Across 12 runs the first position never moved. Positions two to four did. If your plan depends on the exact order of the second and third repair, that ordering is not something we can currently justify.
Our simulation of a complete catalogue can only use values the shelf already records. Where an attribute's recorded values are unrepresentative of its true range, the simulation is wrong in a way we can name but cannot yet correct. It needs a value distribution drawn from outside the catalogue being repaired.
Not yet. When that changes we will say so with their permission, and not before. Every number on this site comes from catalogues we measured ourselves, and each one says which.
Before you commit
Questions
One category export is the cheapest way to find out whether any of this holds for your catalogue.
Request a Decision AuditIf the audit finds nothing useful, we will tell you that on the call.