Free AI Bot & Search Crawler Checker

Verify whether your page grants or restricts access to generative AI training bots like GPTBot, ClaudeBot, PerplexityBot, and search engines like Googlebot and Bingbot.

Check a Page

Quick Test Samples:

Fetching Page Directives & Header Logs...

Verifying User-Agents: GPTBot, ClaudeBot, Googlebot...

Inspected Page Result

AI Agent & Scraper Status

/

GPTBot, ClaudeBot, Perplexity, Cohere, Bytespider

Search Engine Crawler Status

/

Googlebot, Bingbot, YandexBot, DuckDuckBot

Direct AI crawlers

Crawler Owner Purpose robots.txt

Search-mediated crawlers

Crawler Owner Purpose robots.txt

Pages sampled during the crawl

A sample of internal links discovered from the submitted page, not a complete site audit.

Clear current report

About AI Crawlers & Search Engine Bots

Web crawlers visit web pages to extract data for model training, real-time query retrieval, or search engine indexing. Understanding how distinct User-Agents operate helps you control your intellectual property while safeguarding SEO rankings.

AI Training Scrapers

Bots like GPTBot, ClaudeBot, and Google-Extended crawl site content specifically to train foundation LLM models. Blocking them prevents model training without directly harming your Search Engine visibility.

User-Agent Examples: GPTBot, ClaudeBot

Real-Time AI Search

Bots like PerplexityBot and ChatGPT-User perform real-time web lookups to answer user queries with live citations. Disallowing them stops your site from being cited as a direct source in user AI answers.

User-Agent Examples: PerplexityBot

Search Indexers

Bots like Googlebot and Bingbot crawl pages to index content for search engine results pages (SERPs). Blocking these crawlers will remove your website from Google and Bing search results.

User-Agent Examples: Googlebot, Bingbot

How to Use the AI Crawl Checker

Enter your domain and instantly see which AI crawlers can access your site.

1. Enter the Page URL

Paste the exact page you want to audit, including https://, so the checker can inspect the same URL AI crawlers would request.

2. Fetch Crawl Directives

Run the scan to fetch the page response, robots.txt rules, llms.txt signals, and crawler-facing access directives.

3. Review Bot Access

Compare AI training bots, real-time AI search agents, and search indexers to see which User-Agents are allowed or blocked.

Frequently Asked Questions

It fetches a page's raw HTML and robots.txt, then checks whether known AI crawlers - such as GPTBot, ClaudeBot, Google-Extended, and PerplexityBot - are allowed to access it and whether the page's content is present without JavaScript.

Direct AI crawlers used for training or on-demand fetches (OpenAI's GPTBot/OAI-SearchBot/ChatGPT-User, Anthropic's ClaudeBot/Claude-SearchBot/Claude-User, PerplexityBot/Perplexity-User, Common Crawl's CCBot, Amazonbot, Meta-ExternalAgent, Bytespider), plus search-mediated crawlers that feed AI answer engines (Googlebot, Google-Extended, Bingbot, Applebot).

Enter a page URL and click Check crawlability. The page and its robots.txt/llms.txt are fetched server-side and evaluated for each AI crawler.

It means very little text was found in the raw HTML relative to the number of script tags on the page - a sign the main content may be injected by JavaScript after the page loads, which most AI crawlers do not execute. This is a heuristic, not a full browser-rendered comparison.

Some site owners block AI training crawlers to prevent their content from being used to train AI models, while still allowing search engines to index their pages.

Allowing AI crawlers can help your content appear in AI-powered answers and search experiences, such as ChatGPT Search or Perplexity.

No. Robots.txt is a voluntary standard. Reputable AI crawlers respect it, but it does not technically prevent access - for stronger enforcement, use server-level blocking.

If robots.txt is missing, crawlers are generally allowed to access the entire site by default. llms.txt is optional and does not override robots.txt rules or make client-rendered content readable.

The checker follows a handful of internal links found on the submitted page and runs the same lightweight checks on them. This is a small sample for context, not a full site crawl.

Yes. The checker is completely free and does not require sign-up.

Still stuck? Contact us.

Explore More Tools

Discover more free tools to help you build, test, and optimize your site.