Know which AI systems read your content and decide which to allow.

Salience reads your CDN or server logs, checks each AI crawler's identity against the ranges its provider publishes, and separates training crawls from AI search indexing and fetches for a live user's question. You see which pages each system reads and what your site returned. It reports and never blocks.

Free plan, no card, nothing added to your pages. AI crawler detection is on every plan.

The AI Crawlers page for demo-site.example, last 24 hours: training on your content 1,536 requests, indexing for AI search 4,442, answering people about you 331. Crawlers: Amazonbot, ai search, 4,090 requests, Verified; ClaudeBot, training, 1,234 requests, Verified; PerplexityBot, ai search, 304 requests, Verified; ChatGPT User, answering, 268 requests, Verified; ByteSpider, training, 110 requests, 110 Unverified. Bot identity evidence for ChatGPT User: 265 verified requests, 3 unverified claims. For ByteSpider: 0 verified requests, 110 unverified claims. Review the verification evidence before deciding whether to block requests.

The AI systems we classify
  • GPTBot
  • ClaudeBot
  • PerplexityBot
  • Amazonbot
  • Meta AI
  • MistralAI

Sort AI requests by what they are for

Treating every AI request as training misreads most of it. Salience classifies known AI crawlers by purpose, taken from each provider's published documentation, because the right response depends on the purpose.

  • Model training: crawlers collecting content for training, such as GPTBot and ClaudeBot
  • AI search indexing: crawlers building the indexes that AI search engines cite, such as OAI-SearchBot and PerplexityBot
  • User-triggered retrieval: a system fetching a page to answer a live user's question, such as ChatGPT-User
  • Unverified or unknown: requests claiming an AI identity that cannot be verified, and systems Salience has not yet labelled

Check that each AI crawler is who it says it is

An AI user agent is easy to fake. For supported systems, Salience checks the claimed identity against the provider's published IP ranges and reports a verification status per crawler.

  • Claimed crawler identity and user agent
  • Verification status per crawler
  • Source IP and IP-range information
  • Requested paths and status codes
  • Allowed versus rejected response outcomes per bot

Bots and Crawlers on demo-site.example, last 24 hours: 15,541 bot requests, 12,233 verified, 655 failed verification, 2,653 other status. Served versus rejected: Googlebot 357 served, 34 rejected, Impersonators only, 0 verified · 34 unverified; Amazonbot 3,934 served, 16 rejected, The real bot, 16 verified · 0 unverified; Attack Path Probe 413 served, 104 rejected, Not verifiable, 0 verified · 104 unverified.

See which content each AI system reads

Access is recorded per URL, so the question moves from how many requests to which pages, which sections and how often.

  • Access by URL, directory and site section
  • Unique pages per crawler, over time
  • Status codes returned to each system
  • Documentation, articles, pricing, product data, research and paid sections, compared
  • Which content areas each system is reading more of than last period

The bot detail page for ClaudeBot (AI Crawlers) on demo-site.example, last 24 hours: Total Requests 1,234, Unique Paths 673, Status Codes 3, Period 24h. Activity over time in fifteen-minute buckets, peaking at 39 requests. Response status codes: 200 703, 301 527, 404 4. Top paths accessed: /robots.txt 59, /sitemap.xml 20, / 3, /page-98fef0 3, /page-76fe43 3, /page-b63519 3.

What this looks like in practice

ScenarioAI search specialist

I need to know if GPTBot, ClaudeBot and OAI-SearchBot are actually reading my content. Are they? Which pages?

A robots.txt edit that cut ClaudeBot by 40%

  1. Last weekGPTBot fetched 847 product pages, each request verified against OpenAI's published IP ranges.
  2. WednesdayClaudeBot activity dropped 40% in a day.
  3. Same daySalience shows the drop began at the time of a robots.txt update, and the change is reverted.

The per-bot breakdown and the robots.txt change history put the cause and the fix in the same view.

Scenariopublisher, legal and licensing

An AI company wants to license our archive. We have no record of what its crawlers already fetch from us.

An access record for a licensing conversation

  1. The bot report is filtered to the provider's declared crawlers, each classified as model training, AI search indexing or user-triggered retrieval.
  2. Per-section access over the previous 90 days shows the training crawler concentrated on /archive and /research, and the robots.txt audit lists the requests that contradicted the published Disallow rules.
  3. The record is exported with system, URL, timestamp, response and verification status per request, and handed to counsel.

The licensing discussion starts from observed access, with the systems, the sections, the volume and the period taken from the site's own logs.

See which pages AI assistants send people to

A visit is counted as AI-referred when the person arrived from ChatGPT, Perplexity, Claude, Gemini, Copilot or another assistant, or the URL carried a matching source tag. Salience lists the pages those visitors land on, what they open next and how the traffic compares with the previous period, and sets each AI company's fetches against the visitors it sent back.

  • Landing pages per assistant, with the share of human traffic each one accounts for
  • Fetches made by each AI company against the visitors its answers sent back, and fetches per visit
  • Pages that earn impressions in Search that no AI system has fetched, with Search Console connected
  • Sessions, engagement and key events for AI-referred visitors, with Google Analytics 4 connected

Use the request record when you set an AI access policy

Whether your policy is open, licensed or restricted, it needs evidence. From the same log stream, teams can answer:

  • Which AI systems are accessing our website, and at what volume?
  • Are they training, indexing or retrieving?
  • Which commercial or protected sections are they requesting?
  • Are declared crawlers respecting our robots.txt rules?
  • Which crawlers receive successful responses, and which are rejected by our existing infrastructure?
  • Are unknown systems impersonating recognised AI crawlers?
  • Are undeclared scrapers reading the same sections, visible as volume and pattern from one network even when no crawler name is carried?

How different teams use it

One record of machine access supports several decisions at once.

Content and editorial

Which articles, guides and knowledge bases AI systems read, which they never request, and which pages the assistants send readers to.

Legal and licensing

Observed access to ground a licensing conversation: which systems, which sections, what volume, over what period.

AI search and content teams

Training crawls, AI search indexing and user-triggered fetches counted separately, per system and per section.

Security

Systems claiming an AI identity that fail verification, reported with the source IPs before a rule treats them as trusted.

Platform engineering

How existing CDN and WAF rules respond to each AI system: allowed, rejected or rate-limited.

Technical SEO

AI search indexing alongside classic search crawling in the same views, from the same logs.

Why not the tools you already have

Most teams already have a CDN dashboard, an analytics tool and Search Console, and some run a prompt-monitoring platform. Each answers a different question from Salience, and Salience replaces none of them.

CDN AI-crawler dashboards

What it shows
Counts per AI crawler on one CDN's zone, with the option to block them at the edge.
What Salience adds
The same view for any CDN or server, per URL, with a verification status, the purpose of each system and a robots.txt audit against what was requested.

Prompt-monitoring platforms

What it shows
Sample AI answers from the outside to see whether and where you are cited.
What Salience adds
Records machine access to your site from the inside: which systems read which pages and what they were given. Some teams run both.

Google Search Console

What it shows
Googlebot only. No GPTBot, ClaudeBot, PerplexityBot or any other AI system.
What Salience adds
Every AI crawler, verified where the provider allows it, alongside the search crawlers in the same views.

Client-side analytics

What it shows
Human browser sessions. AI crawlers never run the tag.
What Salience adds
Every AI request as the server logged it, whether or not the client executed JavaScript.

Self-built log queries

What it shows
A parser, a user-agent list and IP-range checks that you keep up to date as providers change them.
What Salience adds
Managed ingestion, 200+ named crawler identities, verification and purpose maintained for you, with alerts when a new system appears.

Common questions

What is the difference between an AI training crawler and an AI search crawler?

A training crawler collects content to train models. An AI search crawler builds an index that may power AI-generated answers, closer to classic search indexing. The same company often runs both under different user agents, which is why Salience classifies them separately.

Can Salience tell whether a crawler is genuine?

For supported systems, yes: claimed identities are checked against official provider IP ranges and each crawler carries a verification status. Where a provider publishes no verification data, traffic is reported as unverified rather than guessed at.

Does Salience show which pages an AI system accessed?

Yes. Access is recorded per URL, with unique pages per crawler, status codes and history, so you can see which sections each system reads.

Can Salience prove that our content was used in an AI answer?

No. Salience shows access: which systems requested which content and what your site returned. Whether that content later appears in an answer is not observable from server logs, and Salience does not claim it.

Does Salience block AI crawlers?

No. Salience is observational. It shows how your existing CDN, WAF and rules respond to each system, and gives you the evidence to change those rules, but enforcement happens in your infrastructure.

Does Salience replace a prompt-monitoring platform?

No. Prompt-monitoring tools sample AI answers from the outside. Salience records machine access to your website from the inside, and counts the human visitors who arrive from an AI assistant's answer. They answer different questions and some teams run both.

Can it show which pages AI assistants send people to?

Yes. A human visit is counted as AI-referred when the referrer is an AI assistant such as chatgpt.com, perplexity.ai, claude.ai, gemini.google.com or copilot.microsoft.com, or the URL carries a matching source tag. Salience lists the landing pages per assistant, what those visitors open next, the share of human traffic they represent, and each AI company's fetches against the visitors it sent back. With Google Analytics 4 connected, sessions, engagement and key events for those visitors appear alongside.

Can it show robots.txt violations?

Yes. Salience audits your robots.txt and reports requests from declared crawlers that contradict your published rules.

Where does Google-Extended fit?

Google-Extended is a robots.txt token, not a crawler. Google's AI training uses content fetched by Googlebot, and the token tells Google whether that content may be used for its AI products. It never appears as a request identity in your logs, so Salience reports Googlebot's activity and audits the robots.txt rules you publish for the token.

Does it need JavaScript on my pages?

No. Collection is server-side, from your CDN or web server logs. There is no tracking script or pixel, so an AI crawler's request is captured whether or not it executed JavaScript, and whether or not your analytics tool filtered it out as a bot.

Does it work outside Cloudflare?

Yes. Vercel, AWS CloudFront, Netlify, Kinsta and Shopify (via Cloudflare) are supported, plus Apache and Nginx through an agent, and historical log imports.

What does the free plan include?

One website, 500,000 requests a month, 30 days of history, real-time analytics, bot and AI crawler detection, all 21 site checks and two email alerts. Solo at $19 a month adds the sitemap and Search Console joins, all 16 alerts by email and Slack, log import and six months of history; webhooks and shared dashboards start on Starter at $49. Every plan except Solo has unlimited users, and no card is needed to start.

Trust & data protection

Privacy and data protection

You are the controller

We process only on your instructions. GDPR Art. 28 DPA on every account, nothing to sign.

UK data residency

AWS eu-west-2 (London). Encrypted in transit (TLS 1.2+) and at rest (AES-256).

Server-side collection

No browser tracking script and no client-side pixel.

No sale, no pooling

Your logs are never sold, never used for advertising, never shared between customers. DPA, sub-processor list and security overview available.

Your logs already show what Google and the AI crawlers are doing.

Free plan, no card, nothing added to your pages. AI crawler detection is on every plan.