AI Search Evidence Lab

Turn repeated AI-answer captures into evidence you can audit: exact denominators, citation persistence, topic-specific source patterns, prompt coverage, fact conflicts, URL failures, and bounded intervention results.

Example data — fabricated observations that demonstrate repeated citations, a broken URL, a fact conflict, and a correction.

Bundled methodology and source registries last verified Jul 28, 2026.

Runs entirely in your browser — nothing you paste is uploaded, and data is saved locally only when you explicitly choose the browser-save action. Clear saved removes that browser copy without changing the text currently in the editor. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.

Feedback
Report a bug

Found something broken in Ai Search Evidence Lab? Let us know what happened — this goes straight to a private triage queue, not a public list.

What will be sent
 No tool inputs, uploads, pasted source, complete results, query parameters, or URL fragments are attached automatically. You can edit or remove the selected passage above. Browser and anti-abuse metadata is processed for spam prevention. 
Local data

Saved targets, named lists, and recent check summaries remain only in this browser.

What this adds to AI-search measurement

  • Observation provenance: exact prompt version, surface, model/checkpoint, retrieval mode, sample group, date, and extraction method.
  • Passage-level evidence: distinguish a known cited URL from a citation whose supporting passage was actually captured.
  • Repeatability: persistence and 95% Wilson intervals instead of a single-run visibility label.
  • Source influence: repeated citations grouped by normalized domain, with prompt and surface breadth shown separately.
  • Correction fixtures: human overrides remain attached to the original observation and classifier version.
  • Intervention discipline: before/after panels remain directional unless compatible control evidence exists.

Packet structure

The top-level object accepts observations, corrections, and interventions. Historical records should be append-only. Change a prompt by creating a new version; change a classifier or denominator by creating a new methodology version. Do not rewrite earlier captures.

Required observation fields

id, capturedAt, promptId, promptVersion, sampleGroup, sampleOrdinal, source, provider, surface, retrievalMode, evaluationState, extractorVersion, methodologyVersion, evidenceState

How to use it

  1. Capture repeated answers under the same prompt version, surface, retrieval mode, locale, and methodology.
  2. Include citations and their passages when the source exposes them. Use a passage hash when retention rules prevent storing the text.
  3. Paste or upload the normalized packet and review excluded observations before reading rates.
  4. Investigate fact conflicts and broken URLs manually. A conflict shows disagreement, not which value is correct.
  5. When evaluating an edit, preserve the deployment and recrawl dates and add a comparable control panel when possible.

Limitations

The lab analyzes supplied evidence; it cannot see undisclosed retrieval, internal ranking, cached context, or every answer shown to every user. Citation frequency is not rank, authority, causal influence, traffic, or conversion. A verified crawler request proves that request, not later use in an answer. Human corrections are counted but never silently rewrite raw evidence.

Frequently asked questions

Does this query AI systems for me?

No. It analyzes observation packets you already captured or exported. Keeping acquisition separate prevents an API execution from being mislabeled as a consumer-product result.

What counts in the denominator?

Only evaluated observations with usable evidence. Refusals, provider failures, and not-evaluated records remain visible but do not become zeroes.

How many repeats do I need?

The ai-search-evidence-v1 contract requires at least three compatible observations before applying a stability label and prefers five. Even then, the result is a sample, not a universal rank.

How does the topic-specific source map work?

Add version-matched prompt records with topic labels. The report groups observed citation domains by those supplied labels and preserves citation appearances, distinct prompt breadth, surface breadth, and captured-passage counts. It does not infer topical authority.

Does a before-and-after increase prove my edit caused it?

No. An uncontrolled change is directional evidence. A compatible control panel makes the inference stronger, but the tool still avoids causal language.

Is pasted data uploaded or stored?

Analysis runs in your browser. Nothing is saved unless you explicitly choose Save in this browser; you can clear that local copy at any time.

Next stepQuotability & Entity-Preserving Rewriter — generate the corrected version.

Feature requests for Ai Search Evidence Lab

Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.

Loading…

➕ Request a feature

New requests are reviewed before they appear here.

Where this tool helps

Common use cases

Measure citation persistence across repeated captures

Analyze compatible prompt runs with explicit numerators, denominators, and uncertainty instead of turning one answer into a visibility claim.

Audit which supplied sources recur by topic

Group normalized citation domains by version-matched prompt topics while keeping appearances, prompt breadth, surface breadth, and passage evidence separate.

Find fact conflicts and broken citation URLs

Surface contradictory extracted values and observed 404 or 410 citations for human investigation without declaring which fact is correct or why a URL failed.

Evaluate a content intervention cautiously

Compare compatible before-and-after observations and preserve a directional interpretation unless a suitable control panel supports a stronger inference.

Preserve an auditable evidence packet

Keep prompt versions, capture provenance, corrections, methodology versions, and normalized exports together for repeatable measurement and review.

Watch the full workflow

AI Search Evidence Lab walkthrough

Read the transcript

AI Search Evidence Lab

This beginner walkthrough explains observation packets, denominators, uncertainty, and citation persistence; analyzes the fabricated example; interprets every major result section; demonstrates browser-local storage and export; and shows the limits that keep the report from becoming an unsupported AI-visibility claim.

Step 1

The AI Search Evidence Lab analyzes answer observations that you already captured. It never sends a prompt to an AI product. Keeping capture and analysis separate prevents an API result from being mislabeled as a consumer-product result and lets you preserve exactly where, when, and how each answer was observed.

Step 2

An observation is one captured answer under known conditions. A prompt version freezes the wording being tested. A denominator is the number of eligible observations used to calculate a rate. Citation persistence asks how often the same normalized URL appears across compatible observations. Passage evidence records the supporting text, or a hash when the text cannot be retained.

Step 3

Use the lab to measure recurring citations across repeated captures, see which supplied domains recur by topic, review prompt coverage, find conflicting facts and failed citation URLs, examine a before-and-after content change cautiously, or hand an auditable packet to another analyst. One isolated answer is not enough for these jobs.

Step 4

The versioned method keeps consumer products, APIs, webmaster reports, and logs as different source types. Only evaluated observations with usable evidence enter a rate. The method needs at least three compatible repeats before assigning a stability label and prefers five. Missing or failed acquisition evidence stays visible but does not become a zero.

Step 5

A packet is a JSON file with observations and optional prompt, correction, and intervention records. Treat history as append-only. If prompt wording changes, create a new prompt version. If the classifier or denominator rules change, create a new methodology version. Do not edit old captures to make a later report look cleaner.

Step 6

The large editor contains a fabricated teaching packet: one frozen prompt, three dated observations, a human correction, and a content intervention. Analyze validates and reports it. Restore example resets the editor. Save, Load, and Clear affect only an optional browser copy. Download appears only after a valid analysis.

Step 7

Select Analyze evidence. The browser parses the JSON, validates every observation, calculates the report, and keeps the packet on this device. The status confirms three analyzed observations. The summary then separates eligible count, mention rate, citation rate, conflicts, failed URLs, and corrections.

Step 8

All three observations are eligible. The target was mentioned in two of three, so mention rate is sixty-six point seven percent. At least one citation appeared in all three, so citation rate is one hundred percent. The fixture also contains one fact conflict, one citation observed as a four-oh-four, and one preserved human correction.

Step 9

The report prints every numerator and denominator. It also shows a ninety-five-percent Wilson interval, which is a range that communicates uncertainty from the small sample. Two of three mentions has a wide interval from twenty point eight to ninety-three point nine percent. Three observations are useful evidence, but still far too few for a universal claim.

Step 10

The main example URL appears in all three eligible runs and has supporting passage evidence in all three, so this sample labels it stable. The invented-study URL appears once and has no captured passage, so it is rare with a passage-coverage gap. Tracking parameters are normalized before appearances are counted.

Step 11

Source influence groups normalized domains and separates citation appearances from prompt breadth, surface breadth, and passages captured. Topic rows use only labels supplied by a version-matched prompt record. Prompt coverage keeps the prompt ID, version, eligible count, rates, surface, and model together. Repetition in this packet does not prove topical authority or rank.

Step 12

Two captures give different prices for the same example tool: Free and nine dollars per month. The lab reports a conflict, which means only that the observations disagree. It cannot choose the correct price. Verify the value against a maintained fact source and keep the human correction attached to the original record.

Step 13

The invented-study URL carries an observed four-oh-four status and no passage. Investigate it manually. The URL might have been invented, removed, normalized incorrectly, or moved during a migration. A failed HTTP status is evidence of a URL problem; by itself it is not proof that a model hallucinated.

Step 14

The intervention compares two observations before an example update with one after it. Both sides show a one-hundred-percent citation rate, but the panel says insufficient evidence because one side has fewer than three observations. It also has no control observations, so the change is explicitly not causal evidence. Preserve deployment and recrawl dates when building a stronger test.

Step 15

The evidence-gaps card counts one citation appearance without a supporting passage or passage hash. A known URL and captured support are different evidence states. Corrections are calibration evidence and never silently rewrite the raw capture, so the record preserves both what was first observed and what a reviewer later corrected.

Step 16

Save in this browser stores an optional local copy. Load saved restores that copy after an editor change. Clear saved removes the stored copy but deliberately leaves the current editor text unchanged. None of these actions uploads the packet. Use them only when browser storage is appropriate for the evidence you are handling.

Step 17

Download normalized packet saves the parsed prompts, observations, corrections, and interventions as dated JSON. Keep that normalized handoff beside the immutable original capture packet, methodology version, and any extraction logs. Normalization helps comparison; it should never replace the raw source record.

Step 18

For a real study, freeze the prompt version, surface, retrieval mode, locale, and method; collect compatible repeats; retain citations and passages when allowed; review excluded observations before rates; investigate conflicts and failed URLs; and add a comparable control for interventions. The lab cannot see undisclosed retrieval, internal ranking, cached context, traffic, conversions, or every answer shown to every user.

Keep the raw captures. Bound every conclusion.

Freeze the prompt and capture conditions, repeat compatible runs, retain passages or hashes when allowed, and investigate conflicts and URL failures manually. Preserve corrections and methodology changes without rewriting history, then compare like with like over time.