You are reading the agent-optimized layer of this page: the literal markdown we serve to AI crawlers and assistants, shipped in the page source of every visit. Making sure AI reads the right facts about a company is literally what KnitKnot does.

# How we measure — and how wrong we could be

KnitKnot's measurement methodology: what we measure, the noise floor we published, how we test our own grader, and the claims we refuse to make.

## What we measure

Real buyer questions, run weekly across ChatGPT, Claude, Perplexity, and Gemini. Every answer is graded: do you show up, are you recommended, and is what the AI claims about you actually true. Every claim is checked against dated evidence pages, and every metric drills down to the verbatim AI answer behind it.

## AI answers are noisy — and we measured how much

Ask an AI the same question twice and you get different answers. Any tool in this category will show a score moving; the honest question is whether a movement is real or weather. We measured our own noise floor with a placebo test: we pointed our production impact estimator at a time window where nothing shipped.

  • - 95% null interval on a 15-question cohort: roughly ±23 percentage points of win rate with zero real change.
  • - About one in five null cohorts shows an apparent effect of 10 points or more.
  • - Independent research on AI-answer instability points the same direction (bootstrap confidence intervals of several points on citation share; run-to-run answer overlap sometimes below 1%).

Consequence: we retired every historical per-fix impact number of our own that fell inside the null band, and the placebo test is now a standing build gate.

Full write-up: https://knitknot.ai/experiments/002-placebo-noise-floor/

## We grade our own grader

Every metric we publish is produced by an LLM-judge pipeline, so we test that pipeline against itself: identical answers scored repeatedly with everything pinned, plus gold-labeled fixtures with deliberately planted errors.

  • - Baseline pairwise self-disagreement was up to 16.5% on some stages; we fixed the noisy mechanisms one by one and measured each fix.
  • - The customer-facing head-to-head verdict now reaches 91% pass-pair agreement on identical input, with residual disagreement confined to genuinely hedged answers.
  • - Negative results are published too: several intuitive fixes (worked examples, majority voting on open extraction, bigger judge models) measurably did not help, and we say so.
  • - A committed harness re-runs these measurements on pinned fixtures, so a change that destabilizes the scorer is caught before it ships.

Full write-up: https://knitknot.ai/experiments/001-scorer-repeatability/

## Regression to the mean — the easiest false claim in this category

Content work naturally targets the questions where a brand performs worst, and worst-performing cohorts of a noisy metric rebound on their own: on a window where nothing shipped, our worst-by-win-rate cohorts "gained" +19 percentage points with zero real change. Our rule: any claimed effect on performance-selected questions must beat the measured mean-reversion baseline for that selection rule — not zero.

Full write-up: https://knitknot.ai/experiments/003-selection-bias/

## What we will and won't claim

  • - The AI Presence Score is standing plus trend — a topline thermometer, deliberately not the only thing we report, and never the sole evidence for anything.
  • - Proof of impact is receipts: the changed AI answer citing the page you shipped, with dates — plus a sustained multi-run trend.
  • - We do not publish per-fix point-lift numbers below our own noise floor.
  • - The noise floor is printed next to aggregate movement, not hidden behind it.

## A standing invitation

The full methodology — methods, numbers, honest caveats, and negative results — is an open lab notebook: https://knitknot.ai/experiments/

And for any tool in this category, including us, ask the vendor two questions: what does your estimator report on a window where nothing shipped, and how much do your targeted cohorts improve without intervention? If they can't answer with numbers, their case studies are statistically indistinguishable from weather.

Raw mirror of this content: https://knitknot.ai/methodology.md. Site-wide summary: /llms.txt · full content: /llms-full.txt

How we measure — and how wrong we could be.

Every error bar we publish traces back to an experiment in our open lab notebook: the noise floor of AI answers, the repeatability of our own grader, and the claims we will not make.

What we measure

Real buyer questions, graded every week.

KnitKnot runs the questions your buyers actually ask, weekly, across ChatGPT, Claude, Perplexity, and Gemini — and grades every answer.

Do you show up?

Whether the AI mentions you at all when a buyer asks an open question in your category.

Are you recommended?

Head-to-head: when the AI compares you with competitors, who does it actually pick — and why.

Is it true?

Every material claim the AI makes about you is checked against dated evidence pages, not vibes.

Every metric drills down to the verbatim AI answer behind it. Nothing we report is a number you can't click into.

The noise floor

AI answers are noisy — and we measured how much.

Ask an AI the same question twice and you get different answers. Any tool in this category will show your score moving; the honest question is whether a movement is real or weather. So we ran a placebo test: we pointed our own impact estimator at a time window where nothing shipped. Whatever it reports there is pure noise — and we published it.

±23pp

How far the win rate of a 15-question cohort can swing between two runs with zero real change (the 95% null interval). Roughly one in five null cohorts shows an apparent effect of 10 points or more.

9 → 0

Every one of our own nine historical per-fix impact numbers fell inside the null band. We retired them all, and the placebo test is now a standing build gate on our estimator.

Independent research on AI-answer instability points the same direction: bootstrap confidence intervals of several points on citation share, and run-to-run answer overlap sometimes below 1%. Full method and honest caveats in experiment 002.

Grading the grader

We test our own scoring system against itself.

Every metric we publish is produced by an AI judge. If the judge disagrees with itself, that disagreement lands inside your metrics before the AI engines change anything real. So we feed it identical answers repeatedly, and gold-labeled answers with deliberately planted errors, and measure what comes back.

Repeatability, measured stage by stage

Baseline self-disagreement ran up to 16.5% on some stages. We fixed the noisy mechanisms one at a time and re-measured after each fix — the customer-facing verdict now reaches 91% pass-pair agreement on identical input, with the remainder confined to genuinely hedged answers.

Negative results, published

Several intuitive fixes — worked examples, majority voting, bigger judge models — measurably did not help or made things worse. The write-ups say so.

Enforced going forward

A committed harness re-runs these measurements on pinned fixtures, so a prompt or model change that destabilizes the scorer is caught before it ships.

Customer-facing verdict

91%

Pass-pair agreement on identical input for the head-to-head verdict — the most stable stage we measured, and the one your report is built on.

Method, per-stage numbers, and what each fix did in experiment 001.

Mean reversion

The easiest false claim in this category.

Content work naturally targets the questions where a brand performs worst — and worst-performing cohorts of a noisy metric rebound on their own. On a window where nothing shipped, our worst-performing cohorts “gained” +19 percentage points with zero real change.

Our rule

Any claimed effect on performance-selected questions must beat the measured mean-reversion baseline for that selection rule — not zero. “Your worst questions improved” is never evidence on its own; it is the default expectation under no intervention.

The write-up

Worst, best, random, and real production cohorts replayed through the same null window — including how our issue-based selection compares to raw performance ranking — in experiment 003.

Our standing rules

What we will and won't claim.

The experiments above constrain what we are willing to say — in the product, in sales conversations, and in case studies.

Score = standing + trend

The AI Presence Score is a topline thermometer — deliberately not the only thing we report, and never the sole evidence for anything.

Proof of impact = receipts

The changed AI answer, citing the page you shipped, with dates — plus a sustained multi-run trend. Not a single before/after delta.

No claims below the noise floor

We do not publish per-fix point-lift numbers our own noise floor can’t support. The nine we had are retired.

The floor stays visible

The noise floor is printed next to aggregate movement, not hidden behind it — so you can tell signal from weather too.

A standing invitation

Read the lab notebook. Ask any vendor these two questions.

Every number on this page traces to a public experiment write-up — method, results, honest caveats, and the negative results included.

For any tool in this category — including us — ask:

  1. What does your estimator report on a window where nothing shipped?
  2. How much do your targeted cohorts improve without intervention?

If a vendor can't answer with numbers, their case studies are — statistically speaking — indistinguishable from weather.

The full notebook, experiment by experiment, is at knitknot.ai/experiments.

Start with a benchmark

See how AI represents you — measured honestly.

Start with a free benchmark across ChatGPT, Claude, Perplexity, and Gemini, with every number traceable to the verbatim answer behind it.