AI Crawlers
See which AI crawlers actually visit your site, and how to install tracking.
AI Crawler Tracking shows which AI engines are actually visiting your site: GPTBot, ClaudeBot, PerplexityBot, Bingbot, and more. It complements GEO Audits: audits tell you whether bots can reach a page; crawler tracking tells you whether they do.
Why server-side capture
AI training crawlers do not run JavaScript, so a Google Tag Manager or JavaScript pixel never sees GPTBot or ClaudeBot. Citlyze captures crawler visits server-side, where the real request User-Agent is visible, so even non-JavaScript crawlers are counted.
The human half is different: visitors referred by AI answers do run JavaScript, so for AI Traffic a browser snippet or Tag Manager install works fine (see Install AI Traffic tracking). The server-side installs on this page capture both kinds at once.
Install tracking
Go to Settings → Connections → Website tracking and generate a site key. The key has two parts, a key ID and a signing secret, shown once; copy both. Then add the matching snippet for your platform (for click-by-click walkthroughs of every option, see Install AI crawler tracking):
- Cloudflare (recommended): paste the Worker snippet in front of your site.
- Vercel / reverse proxy: add the middleware snippet.
- WordPress: download AI Crawler Control by Citlyze, upload it in
wp-admin → Plugins, then in its settings fill all three fields: Tracker
base URL (
https://app.citlyze.com), Key ID, and Signing secret. Use Send test event to verify delivery. The plugin does not report until the tracker base URL is filled in. - Custom server / Node: for self-hosted sites (AWS, GCP, Azure, bare
metal, containers): add the Express-style middleware snippet (Node 20+). It
adapts to Fastify, Koa, or plain
http, and other languages can implement the signed-beacon protocol directly.
The Cloudflare, Vercel, and custom server snippets read the signing secret
from an environment variable, CITLYZE_SIGNING_SECRET (an encrypted Worker
secret on Cloudflare), so the secret never sits in code you commit or share.
The WordPress plugin stores it in its settings instead.
Each event is signed with your secret so the tracker can reject forged beacons. That signature proves the report came from your site; it says nothing about who the visitor was. Confirming that a visitor really was GPTBot is a separate step, covered in How verification works.
Standard Shopify stores cannot run server-side capture, and Shopify does not support putting a proxy (such as Cloudflare) in front of a store, so full crawler tracking is not available on standard Shopify. Headless Hydrogen or Oxygen storefronts run server-side code and can use the custom server snippet. Webflow hosting cannot run server code either; on Webflow Enterprise, a self-managed reverse proxy can run the Cloudflare Worker or the custom server snippet at the proxy layer.
Collectors report what happened to each request:
- The Cloudflare Worker, the custom server middleware, and the WordPress plugin send the exact HTTP status (200, 301, 404, 410, 500, and so on), which powers the error and redirect insights below. Vercel middleware runs before the response exists, so it cannot; for status codes on Vercel, use the Vercel log drain under Website tracking → Server logs instead of the middleware.
- For a redirect, the same three collectors also send where it pointed: the path for a page on your own site, only the domain for any other site.
- For crawler visits, the query string is sent too (for example
?page=2). A human visitor's query string never leaves your site.
Events travel in signed batches. If Citlyze is briefly unreachable, the Cloudflare Worker and the custom server middleware retry once, and the WordPress plugin keeps undelivered events in your WordPress database for up to a day and keeps retrying. A retry is never counted twice. Each installed collector also refreshes its list of known crawlers at least once a day, so a newly launched crawler is recognized without updating the snippet or the plugin. The Website tracking page tells you when a newer version of your snippet or plugin is available.
The WordPress plugin sees every request that reaches WordPress, including requests another plugin redirects before the page renders. It cannot see what never reaches WordPress: pages served from a full-page cache, redirects done by your web server or CDN, and requests a firewall blocks before WordPress loads. The plugin's settings page tells you when it detects a page cache. For heavily cached sites, use the Cloudflare Worker or upload your server logs. If WordPress runs behind a proxy or load balancer other than Cloudflare, set the proxy's client IP header and addresses in the plugin settings, so crawlers can be verified by their real address.
Use one collector per site. If a website collector and a server log source both report the same domain, every visit is counted twice; the Install tracking page warns you when that happens.
Custom servers and other languages
The custom server tab ships a Node middleware, but any stack can report: send
signed HTTPS POST requests to https://app.citlyze.com/api/track with a
JSON body (at most 256 KB) and Content-Type: application/json. Each request
carries a batch of 1 to 50 events.
Required headers:
| Header | Value |
|---|---|
x-aeo-schema | The literal string 3. |
x-aeo-key-id | Your key ID (the UUID shown when you generated the site key). |
x-aeo-ts | Unix timestamp in seconds; must be within 5 minutes of the tracker's clock. |
x-aeo-nonce | Unique per batch: 16 to 64 characters of hex digits and dashes. A UUID with the dashes stripped works. |
x-aeo-signature | Lowercase hex HMAC, computed as below. |
Computing the signature:
- Derive the signing key: the lowercase 64-character hex SHA-256 of your
signing secret (the full
ctk_...string). Use that hex string's UTF-8 bytes as the HMAC key; do not hex-decode it. - Build the message:
"3\n" + timestamp + "\n" + nonce + "\n" + bodyDigest, wherebodyDigestis the lowercase hex SHA-256 of the exact body bytes you send. Any re-serialization after signing invalidates the signature. x-aeo-signatureis the lowercase hex HMAC-SHA256 of that message with the signing key from step 1.
Keep the signing secret in an environment variable or your platform's secret store, never in source code.
Body fields:
collector: a short name for your integration, for examplemy-app/1.0(letters, digits,.,_,/, and-, up to 64 characters).events: the list of events. Each event has these fields (required unless noted):occurredAt: when the request happened, in milliseconds since the Unix epoch. Events older than 48 hours are ignored.host: the request's host name without the port, or an empty string. Events for a host that is neither your site key's domain nor one of its subdomains are ignored.userAgent: the visitor's User-Agent, up to 1024 characters.path: the request path with a leading/and no query string or fragment, up to 2048 characters.query(optional): the query string without the?, for crawler requests only.visitorIp: the client IP as seen by your server. Behind a load balancer or reverse proxy, take it from the forwarding header your own proxy sets; this field is what crawler identity verification checks.referrer: an empty string, or anhttps://referrer URL with query and fragment stripped. Only meaningful for human visits arriving from AI answers.utmSource(optional): theutm_sourcevalue, when a visit from an AI answer arrives without a referrer.status: the three-digit HTTP status of the response, orunknownif you report before the response exists.method: the uppercase HTTP method.redirectTarget(optional): for a 3xx response, theLocationit pointed to.
Report only requests whose user agent looks automated or whose referrer is
an AI answer engine, and send batches after the response, never while a
visitor waits. If a batch fails with a network error, a 429, or a 5xx,
retry it with the same nonce and body and a fresh timestamp: a batch
that already arrived is recognized by its nonce and counted once. A 429
carries a Retry-After header. Any other 4xx means the batch itself is
invalid; do not retry it.
Optionally, fetch the current list of known crawler tokens and AI referrer
hosts once a day from GET https://app.citlyze.com/api/track/registry, with
your key ID in the x-aeo-key-id header. The x-aeo-registry-signature
response header is the lowercase hex HMAC-SHA256 of
"citlyze-registry-1\n" + bodyDigest, keyed with the signing key from step
1; ignore the list unless it matches.
To verify your integration, send a batch with "test": true holding a single
event whose userAgent starts with citlyze-connection-test/ and whose
path is /citlyze-test. A correctly signed test returns HTTP 200 and
stores nothing; real batches return 204, whether or not the visits end up in
your reports.
Rotate or revoke a key
If a signing secret may have leaked (for example it was committed to a repo or
shared in a screenshot), use Rotate next to the key. Rotation issues a new
secret for the same key: the key ID, name, domain, and all recorded history are
kept, but the old secret stops verifying immediately, so update
CITLYZE_SIGNING_SECRET (or the plugin's setting) with the new secret right
away. Revoke is different: it disables the key permanently and stops
tracking for that site.
Reading the analytics
The Diagnose → AI crawler activity → Overview page shows, for the selected time frame:
- headline tiles: total crawler hits, distinct crawlers, top crawler, and trend
- crawler visits over time (daily totals)
- when crawlers visit: hits by hour of day
- crawler trends over time (one line per crawler)
- by crawler breakdown with organization, purpose, hits, and trend
- most-crawled pages (top 10, with a link to the full list)
- unrecognized bots: bot-like user agents that matched no known AI crawler, grouped under "Unknown bot" so new crawlers surface early
Days and hours follow your browser's time zone, so a visit at 9 pm in New York counts on that day, not the next. The per-page table uses UTC days.
Trends compare full days only. While today is still in progress it is left out of the comparison, and each full day is compared with the same day one window earlier. A crawler with no visits in the earlier window shows New instead of a percentage, and when tracking started after the earlier window began, no trend is shown yet.
Use the presets (7/30/90 days), the custom date range, and the crawler filter to focus the view. Workspaces tracking more than one site also get a site filter. People who arrived from AI answers are covered on AI Traffic.
How verification works
Any client can put GPTBot in its user agent. So a user-agent string is a
claim, not proof, and Citlyze treats it that way: headline crawler numbers
count only traffic we could confirm independently.
Each visit gets one of three confidence levels:
- Verified: we confirmed the visitor's network identity against something the operator publishes. Only these count toward your crawler totals.
- Probable: corroborating evidence, such as your CDN flagging the request as a known bot, but no independent confirmation.
- Unverified: the user agent named a crawler and nothing contradicted it, but we could not confirm it. Shown separately, never added to your totals.
Verification uses whichever of these the operator supports:
| Method | What it proves |
|---|---|
| Signed request | The request carried a cryptographic signature we checked against the operator's published keys. The strongest evidence available. |
| Published IP range | The source address falls inside a range the operator publishes for its crawlers. |
| Reverse DNS | The source address resolves to the operator's domain, and that name resolves back to the same address. |
| Known IP range | The source address falls inside a fixed range documented by the operator. |
| CDN attestation | Your CDN identified the request as a known bot. Corroborating, not conclusive. |
| User agent only | Nothing but the self-reported name. Always unverified. |
Operators that publish nothing to check against can only ever reach unverified. That is a property of the crawler, not a problem with your setup, and it is why the separate views exist rather than one blended number.
Two things worth knowing:
- Agent traffic is counted separately. Tools people drive themselves, like ChatGPT Agent, appear under agent activity rather than in crawler totals, because one person clicking is not the same signal as a crawler indexing you.
- Occasionally we cannot run a check: an operator's IP list may be temporarily unreachable. Those visits stay unverified rather than being counted, so your totals never include something we did not confirm.
Crawled pages
Diagnose → AI crawler activity → Crawled pages is the full drill-down: every page AI crawlers fetched in the selected window, with search, sorting, and pagination. Each row shows hits, share of all crawler traffic, trend versus the prior window, errors, redirects, the top crawler, and last-seen date; expand a row for the per-crawler split and the exact status codes crawlers received.
Insights appear above the table when relevant:
- Crawled but never cited: pages AI engines fetch but never cite in your tracked answers; candidates for clearer, more citable content.
- Pages returning errors: pages that answered crawlers with 4xx/5xx responses. Broken pages cannot be read or cited.
- Pages that redirect: old URLs crawlers keep requesting and get redirected from. Point your internal links and sitemap at the final URLs.
- Pages fetched for the first time: pages no AI crawler had fetched before this window since tracking started. New content showing up here means crawlers found it.
Status codes and redirects need a collector that reports them: the Cloudflare Worker, the custom server middleware, the WordPress plugin, or server logs.
The table can be downloaded as CSV on plans that include data export.
Crawl log
Diagnose → AI crawler activity → Crawl log lists every request verified AI crawlers and signed agents made to your site, one day at a time and in order, grouped into crawl sessions (one crawler's requests with no gap longer than 30 minutes). For each request you see the time, the page and its query string, the status code, and, for redirects, where the redirect pointed and whether the crawler followed it in the same session. Pages fetched more than once in a session are marked.
The crawl log keeps only verified crawler and signed agent requests, never
other visitors, and never stores an IP address or user agent. Values of
query parameters that could carry credentials or personal data (such as
token or email) are replaced with redacted. How many days you can look
back depends on your plan; see
Plans and limits. A day's log
can be downloaded as CSV on plans that include data export.
Access the data via API and MCP
Crawler visits are available programmatically once tracking is installed:
- REST:
GET /api/v1/crawler-events; see the AI Crawler Events resource. - MCP: the
list_crawler_eventstool; see the MCP tools reference.
Both are read-only and scoped to your workspace. The REST resource supports
filtering by crawler_id, tracked site, and exact path; human visits from
AI answers are exposed at GET /api/v1/ai-referrals.