Skip to content
Tool 02Crawler-access diagnostic

Robots.txt + AI-Bot Tester

Check crawl rules, sitemap signals, and exactly which AI answer engines can access your work—the part most robots testers ignore.

Paste any public URL. We read the live response.

Output / AI

Live results

AI

Ready when you are

Enter a public URL above. Your live diagnostics will appear here without a login or paywall.

What is a robots.txt file?

robots.txt is a plain-text file at the root of your site that tells crawlers which parts of your site they may fetch. Google, Bing and the major AI crawlers all read it before crawling. A broken or overly-strict robots.txt can quietly deindex your whole site; a missing one leaves crawling entirely to guesswork.

Why the AI-crawler check matters

These crawlers each do a different job. GPTBot (OpenAI), ClaudeBot (Anthropic) and CCBot (Common Crawl) gather content for model training; Google-Extended governs Gemini training and does not affect Google Search; PerplexityBot indexes your pages for Perplexity's answers. Most robots testers ignore them — AnswerScope reports whether each is allowed, so you can decide about model training separately from live answers. Note: to be cited in a live AI answer you also need the search crawlers OAI-SearchBot (ChatGPT), Googlebot (Google) and Claude-SearchBot (Claude), which is a separate list. See the full AI crawler reference for what each one controls.

What this tool checks

  • robots.txt is present, parses, and how many disallow rules it has
  • Your sitemap is referenced from robots.txt and is valid XML
  • The page is indexable (no accidental noindex / nofollow)
  • Whether GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot are allowed
  • Whether you publish an llms.txt for AI discovery

Frequently asked questions

01

How do I test my robots.txt file?

Enter your domain above. AnswerScope fetches your live robots.txt, checks that it parses, counts your disallow rules, confirms your sitemap is referenced, and — uniquely — reports whether AI crawlers like GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot are allowed or blocked.

02

How do I block AI crawlers like GPTBot and ClaudeBot?

Add explicit rules to robots.txt, for example "User-agent: GPTBot" followed by "Disallow: /". Each AI company uses its own crawler name — GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google AI), PerplexityBot, and CCBot (Common Crawl). This tool shows you which are currently allowed so you can decide.

03

Should I block AI crawlers?

It depends on your goal. Blocking them protects content from being used to train models, but it can also keep you out of AI answers and assistants — which is increasingly where discovery happens. If you want to be cited in AI answers (AEO), you generally want to allow them.

04

What does "Disallow" actually do?

Disallow tells compliant crawlers not to fetch matching paths. It is a request, not a hard block — well-behaved bots (including Google and the major AI crawlers) respect it, but it does not stop a determined scraper. For hard blocking, use server-side or firewall rules.

The next question

Your page is AI-ready. But does AI actually cite you?

Being crawlable is step one. The AI Visibility Checker shows whether ChatGPT, Google AI Overviews, Perplexity and Gemini name and cite your brand, and who they name instead.

Join the waitlist

Keep going

More free tools

All tools →