See what each AI crawler gets from your pages, then keep it that way
ai-readable is an open-source command line tool and GitHub Action. One command prints the robots.txt verdict for every documented AI crawler, how much of the page exists before JavaScript runs, and nine readability checks scored out of 100. In CI it compares each pull request against a baseline and fails when a page regresses.
npx ai-readable example.com/pricing --renderPer-crawler robots.txt verdicts
Nineteen documented AI crawlers, grouped by what blocking them costs you, with the exact robots.txt line that decided each verdict.
Initial HTML versus rendered
Loads the page in headless Chromium and reports how much of the content only exists after JavaScript, heading by heading.
A gate, not a one-time audit
Commit a baseline for your key pages. Every pull request is checked against it and fails on a regression, with the evidence in a comment.
No API key, nothing sent anywhere
Deterministic checks on the HTML and robots.txt your server serves. The only requests are the page itself, robots.txt and llms.txt.

Check once, then gate every deploy
Run it on a page
The command fetches the page, robots.txt and llms.txt, scores the nine checks, and lists every AI crawler as allowed or blocked for that exact path. Add the render flag to see the JavaScript gap.
Fix it with your agent
The repository is also an Agent Skill. Claude Code, Codex and Cursor run the check, pick a recipe for your framework, apply the smallest fix, and rerun until it is green.
Add the Action
One init command writes the config and a workflow. From then on a search bot newly blocked, a noindex leaked from staging, or content moved behind JavaScript turns the pull request red.
Why the initial HTML is what counts
Most AI retrieval crawlers do not execute JavaScript. A pricing page rendered entirely on the client is a title and an empty container to them, so it cannot be cited no matter how good it looks in a browser. ai-readable compares the HTML a crawler downloads with the page after JavaScript ran and shows the difference, including the headings that only appear after rendering.
Blocking a training bot is not blocking search
robots.txt rules copied from a block-all-AI snippet often disallow OAI-SearchBot and ChatGPT-User along with GPTBot. The first two feed live answers and link reading; the third only feeds model training. The per-crawler table separates the three groups and quotes the rule, so you can opt out of training without disappearing from answers.
Regressions are the real problem
A one-time audit finds a fault once. Sites built and refactored with coding agents reintroduce the same faults every week: a layout change moves the H1 into a client component, a preview setting ships noindex to production. The gate catches those on the pull request, before the deploy, and links the recipe that fixes each one.
What it deliberately does not do
It never sends requests with a bot user agent, because CDNs verify real crawlers by IP address and a spoofed request would describe the CDN, not the crawler. It also never asks an AI engine what it thinks of your brand. A green check means the engines can read the page. Whether they mention and cite you is what Citlyze measures over time.
Readable is the first step. Recommended is the goal.
ai-readable tells you whether AI engines can read a page. Citlyze tells you whether they name and cite you for the questions your buyers ask, measured across ChatGPT, Gemini, Perplexity and Google's AI surfaces with the run count behind every rate.
Yes, on the same checklist: nine checks with the same weights. The two can differ on very large pages, because the hosted grader reads the first 256 KB of HTML and the command line tool reads up to 2 MB.
Does AI recommend you when buyers ask?
Track how ChatGPT, Gemini, Perplexity, DeepSeek, ByteDance, Grok, Google AI Overview & AI Mode, and Bing Copilot mention, cite, and rank your brand. Review tracked answers and the gaps to close.
No credit card required
