Launch
AI Crawler Census
Visit
Example Image

AI Crawler Census

See if ChatGPT & Perplexity can read and cite your site

Visit

A free tool to check whether ChatGPT, Perplexity, and Google's AI answers can read and cite your site. Backed by an original crawler-blocking census of the top 5,000 sites' robots.txt rules (September 2026). No signup required.

Example Image
Example Image
Example Image
Example Image

Features

Checks your site's robots.txt against real AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and more)

Backed by an original census of the top 5,000 sites' robots.txt rules

Shareable result page and an embeddable badge

Includes a free llms.txt validator

Use Cases

Site owners checking whether ChatGPT or Perplexity can cite their content

SEO/GEO practitioners auditing a client's AI-crawler exposure

Publishers deciding whether to allow or block AI training and search crawlers

Comments

very useful for generative engine optimization

Built this after finding that 44% of sites blocking GPTBot also block OAI-SearchBot without realizing it's a different crawler with a different job. Wanted a free, no-signup way for anyone to check their own site in seconds, backed by real census data instead of a guess.

custom-img
Solo founder making mobile apps and saas...

I'll probably try this out. Seems like an important tool in 2026

It would be great if it can give advices of how to improve my website after checking.

The point about 44% of sites blocking GPTBot but also accidentally blocking OAI-SearchBot is really useful. Most site owners don't realise those crawlers do different jobs. Would love to see a suggested robots.txt snippet next to each result, so people can allow AI search citations while still opting out of training. Bookmarking this for client audits.

Useful that it separates training crawlers like GPTBot from search crawlers like PerplexityBot. Does the checker also read meta robots or X-Robots-Tag headers, or only robots.txt? Some sites allow a bot in robots.txt but block it with headers, so that could change the result.

Nice that it's backed by actual census data and not guesses. One edge case I run into: a lot of sites allow AI bots in robots.txt, but Cloudflare's "Block AI bots" setting or a WAF rule returns 403 to GPTBot or PerplexityBot anyway. Does the checker make a real request with each bot's user agent, or only parse robots.txt? A live fetch per crawler would catch those silent blocks, and they're probably more common than robots.txt mistakes now.

Ran it on my site clekara.com: 2/2 core checks passed (OAI-SearchBot and PerplexityBot allowed, homepage readable, llms.txt found). I like that it separates search crawlers from training crawlers, that distinction is easy to miss in robots.txt. Does the llms.txt check also look at llms-full.txt?

custom-img
Desarrollador venezolano 🇻🇪 Creé la Ti...

El dato del 44% me dejó pensando: mucha gente bloquea el crawler de entrenamiento sin saber que también está bloqueando las citas en las respuestas de ChatGPT. Como dueño de un sitio web, esto es oro para decidir qué permitir y qué no. ¿Piensan agregar recomendaciones automáticas de qué reglas ajustar en el robots.txt según el objetivo (visibilidad vs. privacidad)? ¡Buen lanzamiento!

custom-img
Indie developer

Helpful tool. I’ve checked my sites. But is it worth $19 a month for monitoring only? Or I missed something?

Premium Products

Comments

very useful for generative engine optimization

Built this after finding that 44% of sites blocking GPTBot also block OAI-SearchBot without realizing it's a different crawler with a different job. Wanted a free, no-signup way for anyone to check their own site in seconds, backed by real census data instead of a guess.

custom-img
Solo founder making mobile apps and saas...

I'll probably try this out. Seems like an important tool in 2026

It would be great if it can give advices of how to improve my website after checking.

The point about 44% of sites blocking GPTBot but also accidentally blocking OAI-SearchBot is really useful. Most site owners don't realise those crawlers do different jobs. Would love to see a suggested robots.txt snippet next to each result, so people can allow AI search citations while still opting out of training. Bookmarking this for client audits.

Useful that it separates training crawlers like GPTBot from search crawlers like PerplexityBot. Does the checker also read meta robots or X-Robots-Tag headers, or only robots.txt? Some sites allow a bot in robots.txt but block it with headers, so that could change the result.

Nice that it's backed by actual census data and not guesses. One edge case I run into: a lot of sites allow AI bots in robots.txt, but Cloudflare's "Block AI bots" setting or a WAF rule returns 403 to GPTBot or PerplexityBot anyway. Does the checker make a real request with each bot's user agent, or only parse robots.txt? A live fetch per crawler would catch those silent blocks, and they're probably more common than robots.txt mistakes now.

Ran it on my site clekara.com: 2/2 core checks passed (OAI-SearchBot and PerplexityBot allowed, homepage readable, llms.txt found). I like that it separates search crawlers from training crawlers, that distinction is easy to miss in robots.txt. Does the llms.txt check also look at llms-full.txt?

custom-img
Desarrollador venezolano 🇻🇪 Creé la Ti...

El dato del 44% me dejó pensando: mucha gente bloquea el crawler de entrenamiento sin saber que también está bloqueando las citas en las respuestas de ChatGPT. Como dueño de un sitio web, esto es oro para decidir qué permitir y qué no. ¿Piensan agregar recomendaciones automáticas de qué reglas ajustar en el robots.txt según el objetivo (visibilidad vs. privacidad)? ¡Buen lanzamiento!

custom-img
Indie developer

Helpful tool. I’ve checked my sites. But is it worth $19 a month for monitoring only? Or I missed something?

Premium Products