AI Training Scrapers
Bots like GPTBot, ClaudeBot, and Google-Extended crawl site content specifically to train foundation LLM models. Blocking them prevents model training without directly harming your Search Engine visibility.
Verifying User-Agents: GPTBot, ClaudeBot, Googlebot...
Inspected Page Result
AI Agent & Scraper Status
GPTBot, ClaudeBot, Perplexity, Cohere, Bytespider
Search Engine Crawler Status
Googlebot, Bingbot, YandexBot, DuckDuckBot
| Crawler | Owner | Purpose | robots.txt |
|---|---|---|---|
| Crawler | Owner | Purpose | robots.txt |
|---|---|---|---|
A sample of internal links discovered from the submitted page, not a complete site audit.
Web crawlers visit web pages to extract data for model training, real-time query retrieval, or search engine indexing. Understanding how distinct User-Agents operate helps you control your intellectual property while safeguarding SEO rankings.
Bots like GPTBot, ClaudeBot, and Google-Extended crawl site content specifically to train foundation LLM models. Blocking them prevents model training without directly harming your Search Engine visibility.
Bots like PerplexityBot and ChatGPT-User perform real-time web lookups to answer user queries with live citations. Disallowing them stops your site from being cited as a direct source in user AI answers.
Bots like Googlebot and Bingbot crawl pages to index content for search engine results pages (SERPs). Blocking these crawlers will remove your website from Google and Bing search results.
Enter your domain and instantly see which AI crawlers can access your site.
Paste the exact page you want to audit, including https://, so the checker can inspect the same URL AI crawlers would request.
Run the scan to fetch the page response, robots.txt rules, llms.txt signals, and crawler-facing access directives.
Compare AI training bots, real-time AI search agents, and search indexers to see which User-Agents are allowed or blocked.
It fetches a page's raw HTML and robots.txt, then checks whether known AI crawlers - such as GPTBot, ClaudeBot, Google-Extended, and PerplexityBot - are allowed to access it and whether the page's content is present without JavaScript.
Direct AI crawlers used for training or on-demand fetches (OpenAI's GPTBot/OAI-SearchBot/ChatGPT-User, Anthropic's ClaudeBot/Claude-SearchBot/Claude-User, PerplexityBot/Perplexity-User, Common Crawl's CCBot, Amazonbot, Meta-ExternalAgent, Bytespider), plus search-mediated crawlers that feed AI answer engines (Googlebot, Google-Extended, Bingbot, Applebot).
Enter a page URL and click Check crawlability. The page and its robots.txt/llms.txt are fetched server-side and evaluated for each AI crawler.
It means very little text was found in the raw HTML relative to the number of script tags on the page - a sign the main content may be injected by JavaScript after the page loads, which most AI crawlers do not execute. This is a heuristic, not a full browser-rendered comparison.
Some site owners block AI training crawlers to prevent their content from being used to train AI models, while still allowing search engines to index their pages.
Allowing AI crawlers can help your content appear in AI-powered answers and search experiences, such as ChatGPT Search or Perplexity.
No. Robots.txt is a voluntary standard. Reputable AI crawlers respect it, but it does not technically prevent access - for stronger enforcement, use server-level blocking.
If robots.txt is missing, crawlers are generally allowed to access the entire site by default. llms.txt is optional and does not override robots.txt rules or make client-rendered content readable.
The checker follows a handful of internal links found on the submitted page and runs the same lightweight checks on them. This is a small sample for context, not a full site crawl.
Yes. The checker is completely free and does not require sign-up.
Still stuck? Contact us.
Discover more free tools to help you build, test, and optimize your site.
Popular Searches