A free tool to check whether ChatGPT, Perplexity, and Google's AI answers can read and cite your site. Backed by an original crawler-blocking census of the top 5,000 sites' robots.txt rules (September 2026). No signup required.
Checks your site's robots.txt against real AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and more)
Backed by an original census of the top 5,000 sites' robots.txt rules
Shareable result page and an embeddable badge
Includes a free llms.txt validator
Site owners checking whether ChatGPT or Perplexity can cite their content
SEO/GEO practitioners auditing a client's AI-crawler exposure
Publishers deciding whether to allow or block AI training and search crawlers
The point about 44% of sites blocking GPTBot but also accidentally blocking OAI-SearchBot is really useful. Most site owners don't realise those crawlers do different jobs. Would love to see a suggested robots.txt snippet next to each result, so people can allow AI search citations while still opting out of training. Bookmarking this for client audits.
Nice that it's backed by actual census data and not guesses. One edge case I run into: a lot of sites allow AI bots in robots.txt, but Cloudflare's "Block AI bots" setting or a WAF rule returns 403 to GPTBot or PerplexityBot anyway. Does the checker make a real request with each bot's user agent, or only parse robots.txt? A live fetch per crawler would catch those silent blocks, and they're probably more common than robots.txt mistakes now.
El dato del 44% me dejó pensando: mucha gente bloquea el crawler de entrenamiento sin saber que también está bloqueando las citas en las respuestas de ChatGPT. Como dueño de un sitio web, esto es oro para decidir qué permitir y qué no. ¿Piensan agregar recomendaciones automáticas de qué reglas ajustar en el robots.txt según el objetivo (visibilidad vs. privacidad)? ¡Buen lanzamiento!
The point about 44% of sites blocking GPTBot but also accidentally blocking OAI-SearchBot is really useful. Most site owners don't realise those crawlers do different jobs. Would love to see a suggested robots.txt snippet next to each result, so people can allow AI search citations while still opting out of training. Bookmarking this for client audits.
Nice that it's backed by actual census data and not guesses. One edge case I run into: a lot of sites allow AI bots in robots.txt, but Cloudflare's "Block AI bots" setting or a WAF rule returns 403 to GPTBot or PerplexityBot anyway. Does the checker make a real request with each bot's user agent, or only parse robots.txt? A live fetch per crawler would catch those silent blocks, and they're probably more common than robots.txt mistakes now.
El dato del 44% me dejó pensando: mucha gente bloquea el crawler de entrenamiento sin saber que también está bloqueando las citas en las respuestas de ChatGPT. Como dueño de un sitio web, esto es oro para decidir qué permitir y qué no. ¿Piensan agregar recomendaciones automáticas de qué reglas ajustar en el robots.txt según el objetivo (visibilidad vs. privacidad)? ¡Buen lanzamiento!
Find your next favorite product or submit your own. Made by @FalakDigital.
Copyright ©2026. All Rights Reserved