See how fast Google finds each story and which AI systems read it.

Salience reads your CDN or server logs beside Search Console and Google Analytics. It shows when Googlebot first requested each story and how often it comes back, which AI systems fetch which sections and what for, and which stories AI assistants send readers to.

Free plan, no card, nothing added to your pages. Bot and AI detection is on every plan; the Search Console join starts on Solo at $19 a month.

AI discovery gaps for daily.example: 369 of 500 pages with 50+ impressions have never been fetched by an AI search bot in this period. /news/2026/09/rail-deal: 18,825 Google impressions, 216 clicks, 0 AI search fetches; /sport/afl-finals: 10,264 Google impressions, 27 clicks, 1 AI search fetches; /opinion/housing-vote: 9,180 Google impressions, 69 clicks, 0 AI search fetches; /news/2026/09/budget: 8,560 Google impressions, 20 clicks, 0 AI search fetches; /sport/epl-week-5: 6,347 Google impressions, 24 clicks, 0 AI search fetches; /news/2026/09/floods: 5,753 Google impressions, 21 clicks, 0 AI search fetches.

The systems that discover and read publisher content
  • Googlebot
  • GPTBot
  • ClaudeBot
  • PerplexityBot
  • Applebot
  • Meta AI

See when Googlebot first requests each story

Crawl budget is the time Google spends on your site, and for a newsroom that time decides how quickly a story is found. Every request Googlebot makes for a story is in your logs, with the status your origin returned and how long it took. Set against your sitemap, a crawl of your own site and Search Console, the record shows how quickly new journalism is discovered and when visibility follows.

The bot detail page for Googlebot on daily.example, last 24 hours: Total Requests 384, Unique Paths 318, Status Codes 4, Period 24h. Activity over time in fifteen-minute buckets, peaking at 38 requests. Response status codes: 200 328, 403 34, 301 21, 404 1. Top paths accessed: / 14, /robots.txt 7, /sitemap-news.xml 3, /news/2026/09/15/council-vote-on-stadium-delayed 3, /sport/afl-finals-team-news 3, /opinion/housing-targets-editorial 3.

  • First verified Googlebot request per story and per section, and every recrawl after it
  • Articles a site crawl or sitemap knows about that Googlebot has yet to request, listed while the story is still current
  • Status code, response time and recrawl timing for every request Google made
  • Search Console impressions, clicks and the queries that showed each story, read against the crawl record

What this looks like in practice

Scenariopublisher, legal and licensing

An AI company wants to license our archive. We have no record of what its crawlers already fetch from us.

An access record for a licensing conversation

  1. The bot report is filtered to the provider's declared crawlers, each classified as model training, AI search indexing or user-triggered retrieval.
  2. Per-section access over the previous 90 days shows the training crawler concentrated on /archive and /research, and the robots.txt audit lists the requests that contradicted the published Disallow rules.
  3. The record is exported with system, URL, timestamp, response and verification status per request, and handed to counsel.

The licensing discussion starts from observed access, with the systems, the sections, the volume and the period taken from the site's own logs.

ScenarioAI search specialist

I need to know if GPTBot, ClaudeBot and OAI-SearchBot are actually reading my content. Are they? Which pages?

A robots.txt edit that cut ClaudeBot by 40%

  1. Last weekGPTBot fetched 847 product pages, each request verified against OpenAI's published IP ranges.
  2. WednesdayClaudeBot activity dropped 40% in a day.
  3. Same daySalience shows the drop began at the time of a robots.txt update, and the change is reverted.

The per-bot breakdown and the robots.txt change history put the cause and the fix in the same view.

See which stories Google keeps coming back to

Page Importance scores every page from 0 to 10 by how often verified Googlebot and Bingbot return to it over the last 90 days, so the stories and sections Google treats as important are listed from your own logs. With Search Console connected, each URL's index status is crossed with the fetches in your logs, which lists the stories Google has indexed but stopped revisiting and the ones it fetched and declined to index, with the impressions each one still earns.

  • Page Importance per story and per section, for Google and for Bing
  • Stories Google has indexed but is no longer revisiting
  • Stories Google fetched and did not index, with the reason Search Console gives
  • Everything known about one URL in one place: requests, crawler fetches, crawl status, sitemap entry, index status, queries and importance

Use Salience beside Search Console, analytics and your CDN

Each tool a publisher already runs answers part of the question. Search Console reports what Google decided, about two days on. Analytics measures readers whose browsers ran the tag. The CDN dashboard counts crawlers on its own zone. Salience reads the access log all of them sit on and adds the part each leaves out.

  • The request behind every Search Console figure: when Googlebot fetched the story, what it was given and when it came back
  • Crawlers and AI systems that never run an analytics tag, counted from the server's own record
  • Google Analytics sessions, engagement and key events per story next to the request record, with organic and AI-referred sessions counted separately
  • The share of real human page requests your analytics tag never recorded
  • Sections of the publication saved as segments, so every report can be read per desk, per section or per subscriber area
  • Alerts to email, Slack or a webhook, and the same record through exports, the API, the CLI and the MCP server

Catch errors and slow responses while a story is getting traffic

Salience watches the origin's responses while breaking-news or social traffic accelerates and alerts while the window is open.

  • 5xx and 429 responses on the stories gaining audience right now
  • WAF and CDN rejections, and whether they landed on readers or Googlebot
  • Origin response times against their own baseline, with image and redirect failures alongside
  • Alerts to Slack, email or a webhook during the window

The Salience dashboard for daily.example, last 24 hours: Traffic Over Time in fifteen-minute buckets, human and bot requests stacked, with two spikes above 1,800 requests. Site Health for the last 30 minutes: traffic 639 normal, 5xx errors 1 with the baseline still learning, bots 52% normal, crawlers 45 normal, latency 0ms normal, 404s 0.2% critical.

Which AI systems fetch each section, and what for

Each AI system is named, checked against the ranges its provider publishes and sorted by the purpose the provider documents: training crawler, AI search and retrieval, or a fetch triggered by a user's question. The record then shows which stories and sections each one requests.

  • 200+ crawler and agent identities recognised by name
  • Identity verified against official IP ranges where the provider publishes them
  • Training crawlers, AI search and retrieval, and user-triggered fetches counted separately
  • First-seen alerts when a new AI system appears on your site
  • Volumes per system, per section, per day, against your own history

Served versus rejected, from the Bots and Crawlers page for daily.example, last 24 hours, showing the AI bots: 13 mixed bots, 2 rejected outright, 446 rejected requests, 31% of verifiable-bot rejections were impersonators. Amazonbot: 3,934 served, 16 rejected with 403 16, the real bot, 16 verified · 0 unverified, on /news/2026/09/vote 403 1 and /sport/finals 403 1 and 3 more. Claude-User: 1 served, 7 rejected with 403 7, the real bot, 7 verified · 0 unverified, on /robots.txt 403 3 and /opinion/housing 403 1 and 3 more. ByteSpider: 108 served, 2 rejected with 403 2, not verifiable, 0 verified · 2 unverified, on /robots.txt 403 2. ClaudeBot: 1,234 served, 0 rejected, 1,234 verified · 0 unverified.

See which stories AI assistants send readers to

A visit is counted as AI-referred when the reader arrived from ChatGPT, Perplexity, Claude, Gemini, Copilot or another assistant, or carried an AI source tag. Salience lists the stories those readers land on, what they read next and how that traffic compares with the previous period, and sets each AI company's fetches against the readers it sent back.

  • Landing stories per assistant, with the share of human traffic each one accounts for
  • Fetches made by each AI company against the visitors its answers sent back, and fetches per visit
  • Stories that earn impressions in Search that no AI system has fetched
  • Google Analytics sessions and engagement for AI-referred readers, where GA4 is connected

Keep a record of every AI fetch, per URL and per system

For every request from a declared AI crawler, Salience keeps the system, the URL, the time, whether the address verified, what your robots.txt said for that path and the status your infrastructure returned. Sections of the publication saved as segments give the same record per desk or per subscriber area.

  • Fetch counts by system and by section, with history and baselines
  • Requests to paths your robots.txt disallows for that crawler, listed with the system, the path and the time
  • Subscriber and premium sections read separately, as segments
  • Exports by CSV, the API, the CLI and the MCP server for commercial and legal teams

What each AI company gives back, from the AI Crawlers page for daily.example, last 24 hours. 0 visitors sent by AI answers · 4,634 search & answer fetches: Amazon (Amazonbot) 3,950 search and answer fetches, 0 training crawls, 0 visitors sent, nothing back; Anthropic (Claude-User, ClaudeBot) 8 search and answer fetches, 1,228 training crawls, 0 visitors sent, nothing back; OpenAI (ChatGPT User, GPTBot, OpenAI SearchBot) 309 search and answer fetches, 1 training crawls, 0 visitors sent, nothing back; Perplexity (PerplexityBot) 304 search and answer fetches, 0 training crawls, 0 visitors sent, nothing back; Common Crawl (CCBot) 0 search and answer fetches, 163 training crawls, 0 visitors sent, —; ByteDance (ByteSpider) 0 search and answer fetches, 110 training crawls, 0 visitors sent, —; DuckDuckGo (DuckAssistBot) 55 search and answer fetches, 0 training crawls, 0 visitors sent, nothing back.

What the record shows in a rights or licensing conversation

The DSM Copyright Directive, the EU AI Act and the GPAI Code of Practice each refer to machine-readable rights reservations, and robots.txt is where most publishers state one. Your server log is the first-party record of whether declared crawlers respected it. Salience keeps that record per request and reports the fetches that contradicted the rule you published.

  • A first-party record of every fetch: system, URL, time, verification result, the rule in force and the response
  • Requests from a declared crawler to a path your robots.txt reserved, documented as they happened
  • Fetch volumes per system before and after a reservation, a block or an agreement
  • Exportable for your counsel, who decide whether and how to use it

The Log Explorer for daily.example filtered to verified bots, bot name ClaudeBot and paths containing /investigations/, last 24 hours. Rows: 2026-09-15 21:42:08 GET /investigations/water-deal/2 answered 200 to ClaudeBot from 160.79.104.21, US, 41.2 KB; 2026-09-15 21:41:55 GET /investigations/water-deal answered 200 to ClaudeBot from 160.79.104.21, US, 38.6 KB; 2026-09-15 19:07:31 GET /investigations/ports-tender answered 200 to ClaudeBot from 160.79.104.18, US, 27.9 KB; 2026-09-15 16:23:14 GET /investigations/ answered 200 to ClaudeBot from 160.79.104.18, US, 22.4 KB; 2026-09-15 13:50:02 GET /investigations/rail-costs answered 200 to ClaudeBot from 160.79.104.21, US, 36.1 KB; 2026-09-15 13:49:47 GET /investigations/water-deal answered 200 to ClaudeBot from 160.79.104.21, US, 38.6 KB; 2026-09-15 11:16:29 GET /investigations/care-homes answered 200 to ClaudeBot from 160.79.104.18, US, 31.7 KB.

See whether access changed after a rule change or agreement

When you change robots.txt, add a block at your CDN or sign an agreement, Salience shows whether the traffic changed: which systems adjusted, which carried on, and what the block stopped.

  • Request volumes per system before and after the change, on the timeline as a site event
  • Restricted paths still being requested after the change
  • Alerts when a system's behaviour changes against its own baseline

The Alerts page for daily.example after a robots.txt change: 2 open · 8 cleared · 10 alerts. Crawler Frequency Change, Warning: CCBot requests to /archive/2019/03/markets-daily rose to 839 after the robots.txt change at 12:00. CCBot, 1 occurrence, 2h ago, first seen 16/09/2026, 14:04, last seen 16/09/2026, 14:04. Crawler Frequency Change, Info: GPTBot requests to /archive/* fell from 1,341 to 0 after the robots.txt change at 12:00. GPTBot, 1 occurrence, 2h ago, first seen 16/09/2026, 14:04, last seen 16/09/2026, 14:04.

How different teams use it

The same request stream answers a different question for each team.

Head of Audience / SEO

How quickly does Google find new journalism, which stories does it keep returning to, and which has it stopped revisiting?

Editorial operations

Which stories are being discovered, recrawled and gaining audience right now, and which ones AI assistants are sending readers to?

Platform engineering

Are errors, response times or access controls affecting an important traffic window?

Rights and licensing

Which AI systems fetch which parts of the publication, how many readers each sends back, and how does access change over time?

Legal and policy

What happened before and after an agreement, a rights reservation or a technical block, from the server's own record?

Why not the tools you already have

Most publishers already run Search Console, an analytics tag and a CDN dashboard, and some use a prompt-monitoring platform or a licensing marketplace. Each answers a different question from Salience.

Google Search Console

What it shows
Search, News and Discover performance, index coverage and crawl statistics, aggregated and reported about two days on.
What Salience adds
The request behind each figure, live: when Googlebot fetched the story after publish, the status it was given and when it came back, with the Search Console figures read against that record and each URL's index status crossed with the fetches in your logs.

Client-side analytics

What it shows
Sessions, pages and events from the tag in the browser, with known bots filtered out by design.
What Salience adds
Every request the server logged, reader or crawler, whether or not the client ran the tag, with the GA4 figures for each story alongside and the share the tag missed.

CDN AI-crawler dashboards

What it shows
Counts per AI crawler on one CDN's zone, with the option to block at the edge.
What Salience adds
The same view for any CDN or server, per section, with a verification status, the published purpose of each system, the requests that contradicted robots.txt and the readers each assistant sent back.

Prompt-monitoring platforms

What it shows
Sample AI answers from the outside to see whether and where you are cited.
What Salience adds
Records which systems fetched which stories from the inside, what your site returned, and which stories the assistants then sent readers to. Some publishers run both.

Licensing marketplaces

What it shows
List content and set terms for the AI buyers that join the marketplace.
What Salience adds
The first-party record of what each declared crawler already fetches, before and after an agreement, exportable for the conversation.

Common questions

How quickly can we see new publishing activity?

Server requests appear as they happen. Search Console's recent performance data is synced daily and runs two to three days behind Google, and it is then read against the first-party request timeline.

Can Salience tell training crawlers from AI search?

Yes, by each system's published purpose. Training crawlers such as GPTBot and CCBot, AI search and retrieval crawlers such as OAI-SearchBot and PerplexityBot, and user-triggered fetches such as ChatGPT-User are counted separately. Where a provider publishes no purpose, the system is listed under the name it declares.

Can it show which stories AI assistants send readers to?

Yes. A human visit is counted as AI-referred when the referrer is an AI assistant such as chatgpt.com, perplexity.ai, claude.ai, gemini.google.com or copilot.microsoft.com, or the URL carries a matching source tag. Salience lists the landing stories per assistant, what those readers open next, the share of human traffic they represent, and each AI company's fetches against the visitors it sent back. With Google Analytics 4 connected, the sessions and engagement for those readers appear alongside.

What does the access record cover?

For every request from a declared crawler: the system, the URL, the time, the verification result, what your robots.txt said for that path and the status your infrastructure returned. That record shows which content was fetched, by whom and when. What a provider did with the content after the fetch happens outside your infrastructure, so the record is presented as access, with the provider's published purpose alongside it.

Where does blocking happen?

At your CDN, WAF or robots.txt. Salience records what each system fetched before a block and whether the requests stopped after it, so the evidence for a rule and the check on its effect come from the same record.

Is the record usable in a licensing or legal conversation?

It is your own first-party access log, classified and exportable: system, URL, timestamp, response, verification result and the robots.txt rule that applied. Whether and how to use it is a question for your counsel.

What about crawlers that ignore robots.txt?

They appear in the log like everything else. Salience fetches your live robots.txt, parses the rule groups and reports every request from a declared crawler to a path its rules disallow, with the system, the path and the time.

Where does Google-Extended fit?

Google-Extended is a robots.txt token. Google fetches publisher content with Googlebot, and the token tells Google whether that content may be used for its AI products. Salience reports Googlebot's requests and audits the rules you publish for the token.

New AI crawlers appear every month. How do you keep up?

Salience recognises 200+ named crawler and agent identities and alerts you the first time a new system appears on your site, so the inventory grows with the ecosystem.

How does the record relate to the EU AI Act?

The AI Act's copyright obligations apply to general-purpose AI providers, and they refer to machine-readable rights reservations such as robots.txt. Salience gives publishers and their advisers the first-party technical record of which declared crawlers fetched what, and whether those fetches respected the reservation you published, to use in that assessment.

Can we read every report by section of the publication?

Yes. Segments are saved page groups defined by path prefix, pattern or query string, with exclusions and nesting, and a one-click library adds common groups such as articles, pagination, search results and parameter URLs. Segments apply to all history, and every report, alert and export can be read per segment.

Which fields does Salience collect, and where is the data held?

Access-log fields: path, method, status, IP, user agent and timing. Data is held in AWS eu-west-2 (London), encrypted in transit and at rest, processed under a GDPR Art. 28 DPA and never shared between customers.

What does the free plan include?

One website, 500,000 requests a month, 30 days of history, real-time analytics, bot and AI crawler detection, all 21 site checks and two email alerts. Solo at $19 a month adds the sitemap and Search Console joins, all 16 alerts by email and Slack, log import and six months of history; webhooks and shared dashboards start on Starter at $49. Every plan except Solo has unlimited users, and no card is needed to start.

Trust & data protection

Privacy and data protection

You are the controller

We process only on your instructions. GDPR Art. 28 DPA on every account, nothing to sign.

UK data residency

AWS eu-west-2 (London). Encrypted in transit (TLS 1.2+) and at rest (AES-256).

Server-side collection

No browser tracking script and no client-side pixel.

No sale, no pooling

Your logs are never sold, never used for advertising, never shared between customers. DPA, sub-processor list and security overview available.

Your logs already show which systems fetch your journalism.

Free plan, no card, nothing added to your pages. Bot and AI detection is on every plan; the Search Console join starts on Solo at $19 a month.