Bot Traffic · August 19, 2026

What's Actually in Your Server Logs? The AI Bot Mix Has Changed Completely

AI bots now account for 22% of verified bot traffic — but over half are training crawlers that send almost nothing back. Here's what the numbers say about who's actually hitting your site.

By the Wrenda team · This article was generated with AI. Figures are sourced where cited below.

Here's a number worth sitting with: bots now account for 53% of all web traffic. Bots outnumber humans on the open web — and that's before you factor in the surge of AI-specific crawlers that have arrived since 2023. Within the verified bot population, AI crawlers have gone from a rounding error to roughly 22% of all bot traffic in under two years. So what does that actually look like in your server logs? And — more importantly — are any of these bots actually sending you meaningful traffic back?

Where does this data come from?

The figures below draw on two publicly released analyses from a major CDN provider (covering bot trends through mid-2026), an independent crawler volume study tracking GPTBot and peers from May 2024 through mid-2025, HUMAN Security's 2026 State of AI Traffic report, and a July 2026 B2B panel study tracking how many pages each AI crawler consumes versus how many visits its parent platform sends back. These are aggregated numbers from real traffic, not surveys or modelling.

So which bots are actually knocking?

Share of AI Bot Traffic by Platform
Percentage share within the AI bot category, CDN network analysis, May–July 2026

Among the named AI crawlers, Meta-ExternalAgent takes the biggest slice at 16.7% of AI bot traffic, followed by GPTBot at 12% and ClaudeBot at around 11.7%. PerplexityBot sits well below those three by volume — under 0.5% — but its growth trajectory is almost vertical. Between May 2024 and May 2025, PerplexityBot's raw request count grew by 157,490%. That's not a typo. It's climbing from a near-zero baseline, but the direction of travel is clear.

GPTBot's own growth is similarly steep: a 305% year-over-year increase in that same May 2024-to-2025 window, with its share of total verified bot traffic jumping from 4.7% to 11.7% between July 2024 and July 2025. ClaudeBot grew from 6% to nearly 10% over the same period. These aren't gradual drifts — they're step changes that outpace most teams' ability to update their analytics or access-control rules.

Are they training bots or retrieval bots? Does it matter?

More than most people realise.

AI Crawler Requests by Purpose
Classification of AI bot requests into purpose categories, May 2026

A May 2026 bot classification breakdown found that 51.8% of AI crawler requests were training-only — meaning the bot is harvesting content to improve a model, not to route any user toward your site. Another 35.7% were mixed-purpose (training and live retrieval), and just 9.3% were purely search and retrieval bots responding to actual user queries.

Put plainly: the majority of what's hitting your servers under an AI user-agent is building someone else's model. The bots that might actually send you a visitor — the ones triggered when a user asks an AI assistant "recommend a good X" and the system fetches relevant pages — are a minority of the AI crawl load.

This has practical implications for access control. A training bot sweeping your entire sitemap overnight consumes real bandwidth and server resources. A retrieval bot that fetches your pricing page in response to a live query has a plausible path to turning into a customer. They warrant different treatments — and most robots.txt configurations currently don't distinguish between them.

What's the crawl-to-refer gap, and why does it matter?

Pages Crawled per Referred Visit (July 2026)
Crawl-to-refer ratio: pages consumed per single referral sent back. Lower is better.

Think of the crawl-to-refer ratio as: for every page an AI platform's crawler takes from your site, how many times does its platform actually send a human visitor back to you? For Google's crawler, that ratio sits at roughly 5:1. For GPTBot, the July 2026 figure is 217:1. For ClaudeBot it's 2,237:1 — down from over 3,300:1 in June as the underlying platform scaled its citation features. For MistralBot the figure sits at 3,389:1.

These ratios are moving in the right direction. But the orders-of-magnitude gap compared to traditional search crawlers hasn't closed. For training crawlers specifically, there's arguably no referral path at all — they're collecting datasets, not routing queries.

To put the opportunity in perspective: Adobe measured a 393% year-over-year jump in AI-referred visits to US retailers in Q1 2026. ChatGPT drove 89% of all AI referrals as recently as October 2025, though that share has diversified to around 63% by mid-2026 as other AI assistants matured. AI search visits globally grew 42.8% year-on-year between Q1 2025 and Q1 2026 — from 15.6 billion to 27.4 billion sessions. The aggregate referral market is growing fast. The challenge is that individual sites are bearing a crawl load that isn't proportional to what comes back yet.

What should you actually do with this?

Start with a log audit. Pull your server logs, filter for known AI bot user-agent strings, and see what's actually there. Most teams are surprised both by the volume and by the specific mix — which bots dominate varies significantly by industry and content type.

Once you know who's showing up, think about differentiating between them. Training-only crawlers with crawl-to-refer ratios measured in the thousands are a different proposition to retrieval bots plausibly connected to live user queries. Reverse proxies, CDN rules, and robots.txt directives can treat these separately if you know the UA strings — and the growing number of AI platforms that publish explicit documentation of their crawl behavior makes this increasingly practical.

Finally, establish a monitoring baseline now. The bot-mix data that's accurate today can look completely different twelve months later. PerplexityBot's trajectory is the clearest example: virtually absent in mid-2024, growing at a rate no quarterly review cycle can keep pace with. A lightweight user-agent segment in your analytics stack, checked monthly, will catch shifts before they turn into surprises.

Sources

  1. From Googlebot to GPTBot: Who's crawling your site in 2025
  2. The crawl-to-click gap: AI bots, training, and referrals
  3. AI Crawler Volume Growth 2022-2026: GPTBot +305% YoY
  4. 2026 State of AI Traffic and Cyberthreat Benchmark Report
  5. Crawl-to-Refer Ratio 2026: AI Bots Take 4,580, Give 1