Start a 14-day free trial with no credit card

Which AI crawlers can actually read your site?

Enter a URL and we read your robots.txt the way crawlers do, then report the verdict for 44 documented AI crawlers: allowed or blocked, the exact rule that decided it, and what blocking each category costs you. No account, nothing stored.

how it works

One URL in, a verdict per crawler out

01

Enter any page URL

We resolve redirects first, so the check runs against the origin that actually serves your site, then fetch its robots.txt exactly once.

02

We evaluate every documented AI crawler

Each crawler's robots tokens are matched against your rules under RFC 9309 semantics: longest match wins, wildcards included. The winning rule is shown next to every verdict.

03

Read the result by category

Search-index and assistant-fetch crawlers gate whether AI answers can see and cite you; training crawlers only affect future model training. The tool keeps the two decisions separate.

What a robots.txt block does and does not do

robots.txt is the documented contract between your site and well-behaved crawlers. Blocking a retrieval crawler removes your pages from the answers that engine builds; blocking a training crawler only opts your content out of future model training. The file cannot stop a bad actor, and it does not need to: the crawlers that matter for AI visibility all publish their tokens and honor the rules.

Why we do not probe your site with bot user agents

Sending requests that pretend to be GPTBot would test how your site treats fake GPTBot, which is a different thing: CDNs verify real crawlers by their published IP ranges, not by the user-agent string. A spoofed probe would produce confident, wrong answers. Reading robots.txt reports the contract you actually publish.

The block nobody remembers writing

The most common cause of zero AI visibility is not content quality. It is a blanket bot rule added years ago at the CDN or in robots.txt, often during a scraping scare, that still catches every new AI crawler today. If this tool shows retrieval crawlers blocked and you do not remember deciding that, you have found your highest-leverage fix.

CDN and firewall blocks will not show here

This tool reads robots.txt, the layer you publish. Some sites also block bots at the CDN or firewall layer, which is invisible from the outside and can contradict a permissive robots.txt. If your robots.txt allows a crawler and its operator still cannot fetch you, check your CDN's bot management settings next.

the next question

Access is step zero. Visibility is the measurement.

This checker tells you whether AI crawlers may read your site. Citlyze tells you what AI engines actually say about your brand once they have: mentions, citations, and competitor share of voice, tracked on a schedule.

FAQ

Crawler access, answered

Want a demo of Citlyze? Book a demo

GPTBot collects content for model training; OAI-SearchBot indexes for ChatGPT's search answers, and a separate user agent fetches pages a user asks about live. Blocking the training bot leaves your search visibility alone; blocking the search and fetch agents removes you from ChatGPT answers. The same split exists at other vendors, which is why this tool groups verdicts by category.

Does AI recommend you when buyers ask?

Track how ChatGPT, Gemini, Perplexity, DeepSeek, ByteDance, and Grok mention, cite, and rank your brand. Review tracked answers and the gaps to close.

No credit card required