How we measure — and how wrong we could be.
Every error bar we publish traces back to an experiment in our open lab notebook: the noise floor of AI answers, the repeatability of our own grader, and the claims we will not make.
What we measure
Real buyer questions, graded every week.
KnitKnot runs the questions your buyers actually ask, weekly, across ChatGPT, Claude, Perplexity, and Gemini — and grades every answer.
Do you show up?
Whether the AI mentions you at all when a buyer asks an open question in your category.
Are you recommended?
Head-to-head: when the AI compares you with competitors, who does it actually pick — and why.
Is it true?
Every material claim the AI makes about you is checked against dated evidence pages, not vibes.
Every metric drills down to the verbatim AI answer behind it. Nothing we report is a number you can't click into.
The noise floor
AI answers are noisy — and we measured how much.
Ask an AI the same question twice and you get different answers. Any tool in this category will show your score moving; the honest question is whether a movement is real or weather. So we ran a placebo test: we pointed our own impact estimator at a time window where nothing shipped. Whatever it reports there is pure noise — and we published it.
How far the win rate of a 15-question cohort can swing between two runs with zero real change (the 95% null interval). Roughly one in five null cohorts shows an apparent effect of 10 points or more.
Every one of our own nine historical per-fix impact numbers fell inside the null band. We retired them all, and the placebo test is now a standing build gate on our estimator.
Independent research on AI-answer instability points the same direction: bootstrap confidence intervals of several points on citation share, and run-to-run answer overlap sometimes below 1%. Full method and honest caveats in experiment 002.
Grading the grader
We test our own scoring system against itself.
Every metric we publish is produced by an AI judge. If the judge disagrees with itself, that disagreement lands inside your metrics before the AI engines change anything real. So we feed it identical answers repeatedly, and gold-labeled answers with deliberately planted errors, and measure what comes back.
Repeatability, measured stage by stage
Baseline self-disagreement ran up to 16.5% on some stages. We fixed the noisy mechanisms one at a time and re-measured after each fix — the customer-facing verdict now reaches 91% pass-pair agreement on identical input, with the remainder confined to genuinely hedged answers.
Negative results, published
Several intuitive fixes — worked examples, majority voting, bigger judge models — measurably did not help or made things worse. The write-ups say so.
Enforced going forward
A committed harness re-runs these measurements on pinned fixtures, so a prompt or model change that destabilizes the scorer is caught before it ships.
Customer-facing verdict
91%Pass-pair agreement on identical input for the head-to-head verdict — the most stable stage we measured, and the one your report is built on.
Method, per-stage numbers, and what each fix did in experiment 001.
Mean reversion
The easiest false claim in this category.
Content work naturally targets the questions where a brand performs worst — and worst-performing cohorts of a noisy metric rebound on their own. On a window where nothing shipped, our worst-performing cohorts “gained” +19 percentage points with zero real change.
Our rule
Any claimed effect on performance-selected questions must beat the measured mean-reversion baseline for that selection rule — not zero. “Your worst questions improved” is never evidence on its own; it is the default expectation under no intervention.
The write-up
Worst, best, random, and real production cohorts replayed through the same null window — including how our issue-based selection compares to raw performance ranking — in experiment 003.
Our standing rules
What we will and won't claim.
The experiments above constrain what we are willing to say — in the product, in sales conversations, and in case studies.
Score = standing + trend
The AI Presence Score is a topline thermometer — deliberately not the only thing we report, and never the sole evidence for anything.
Proof of impact = receipts
The changed AI answer, citing the page you shipped, with dates — plus a sustained multi-run trend. Not a single before/after delta.
No claims below the noise floor
We do not publish per-fix point-lift numbers our own noise floor can’t support. The nine we had are retired.
The floor stays visible
The noise floor is printed next to aggregate movement, not hidden behind it — so you can tell signal from weather too.
A standing invitation
Read the lab notebook. Ask any vendor these two questions.
Every number on this page traces to a public experiment write-up — method, results, honest caveats, and the negative results included.
For any tool in this category — including us — ask:
- What does your estimator report on a window where nothing shipped?
- How much do your targeted cohorts improve without intervention?
If a vendor can't answer with numbers, their case studies are — statistically speaking — indistinguishable from weather.
The full notebook, experiment by experiment, is at knitknot.ai/experiments.