AI bots by intent, and what each AI company gives back

AI traffic is now split into training, AI search and user-triggered fetches, robots.txt violations ignore fetchers, and a new report shows fetches taken versus visitors sent back per AI operator plus the pages AI search has never seen.

Three kinds of AI visitor. Every AI bot now carries an intent: training (corpus crawls such as GPTBot and ClaudeBot), search (index crawls such as OAI-SearchBot and PerplexityBot) and fetcher (a real person just asked the AI about your page: ChatGPT-User, Claude-User, Perplexity-User). The LLM Crawlers page shows the split as three tiles you can click to filter, every crawler carries a badge, and the same field flows through the API, MCP, CLI and exports.

Fetchers are not robots.txt violations. User-triggered fetchers ignore robots.txt by design, like a browser. The robots.txt report now lists them separately instead of counting them as violations.

What each AI company gives back. For each operator: how many pages its search and answer bots fetched, how many training crawls it made, how many people then arrived on your site from its answers, and the ratio between them. Visitors are detected from the referrer or from the utm_source=chatgpt.com-style tags AI products add to their links, and those visits are now always kept, even if you store bots only.

AI discovery gaps. Pages Google shows often (Search Console impressions) that no AI search bot has fetched in the period, invisible to AI search until they are.

← A living bot registry, and name your own botsThe API, MCP and CLI can now act, and recommendations come with ready-to-paste rules →
← All changelog entries