Bot Traffic · September 9, 2026

Which AI Bots Are Actually Hitting Your Site — and Do Any of Them Send You Traffic?

AI-related requests now make up over a quarter of all verified bot traffic on major networks. But almost none of that actually sends referral clicks. Here's what the 2026 data says about what's crawling your site and why.

By the Wrenda team · This article was generated with AI. Figures are sourced where cited below.

If you pull your server logs this week, there's a good chance AI-related bot requests make up over 20% of your verified bot traffic. That figure crossed 26% of all bot requests on major content delivery networks by May 2026, up from near zero four years ago. The uncomfortable follow-up question: how much of that ever sends you a visitor?

The short answer is very little. The crawlers doing the most crawling aren't the ones sending you traffic. And if you haven't looked at your bot traffic breakdown recently, you probably don't know which category most of yours falls into.

Where does this data come from?

The figures in this post are drawn from network-level analyses covering tens of billions of monthly HTTP requests across mid-2026, robots.txt audits of the Tranco top-10,000 domains, and referral attribution studies from analytics platforms tracking millions of sites. Bot traffic measurement isn't perfectly consistent across providers — different collection points, different filtering thresholds — so we've cross-referenced multiple sources throughout.

So who's actually crawling?

Top AI Crawlers by Traffic Share (July 2026)
Share of total AI bot traffic across major CDN networks, July 2026. ClaudeBot led all 31 days, finishing at 16.28%; GPTBot held steady at 9.74%.

The AI crawler landscape in 2026 is more fragmented than most people assume. Yes, GPTBot and ClaudeBot get coverage in tech press, but the actual traffic picture also includes Meta-ExternalAgent, CCBot (the Common Crawl scraper that has been around since 2012), Applebot, OAI-SearchBot, and PerplexityBot, among others.

In July 2026, ClaudeBot accounted for roughly 16% of all AI bot traffic and GPTBot around 10%. But the total AI bot universe is bigger than just those two — Meta-ExternalAgent, CCBot, and Applebot all have significant share. And the composition shifts month to month: ClaudeBot hit 19.8% of AI bot traffic in June 2026 before coming back down. GPTBot has been fairly steady at 9 to 12% across the year.

What's easy to miss here: these bots don't all behave the same way. Some run large batch indexing sweeps over days — you'll see a surge in requests across your full sitemap, then nothing for weeks. Others are real-time agents responding to actual user queries, following links on-the-fly in a way that looks much more like a browser session than a crawler. They arrive under similar user-agent strings. What they're doing with your content is completely different.

What are they actually after?

AI Crawler Requests by Declared Purpose (August 2026)
Training crawlers accounted for ~40% of AI bot requests while search-linked crawlers made up 16%. The user-action bucket nearly doubled vs H1 2026 as agent browsing grows.

This is where the picture gets genuinely surprising. Of all AI crawler requests in August 2026, roughly 40% were explicitly for training purposes — building or updating a model's knowledge base. Only around 16% were linked to search and answer features — the kind that might actually surface your content to a user. The rest falls into mixed-use categories, undeclared, or the fast-growing user action bucket, which covers agents browsing on behalf of a real person.

That 40 to 16 split matters more than most site operators realise. A training crawler ingests your content and bakes it into parameters somewhere. It won't send you referral traffic. It might not even index you in a way you can verify or influence. A search-mode crawler, by contrast, is what could eventually serve your page when someone asks an AI assistant a relevant question. These two categories share IP ranges and can look nearly identical in a raw access log.

The practical question for most sites isn't whether they're being crawled by AI — they almost certainly are. The question is which kind.

Are publishers trying to separate them?

The blocking picture has shifted significantly. Across the general web, about 8 to 10% of sites block GPTBot via robots.txt. Among news publishers and content-heavy sites, that figure exceeds 50%. Publishers are blocking AI training crawlers at five to seven times the rate of the rest of the web, and the gap has been widening through 2026.

The more nuanced approach that has emerged — used by around 30% of top sites — is selective blocking. They disallow training-focused crawlers while explicitly permitting search-mode bots like OAI-SearchBot and PerplexityBot. The bet: stop giving away training data for free, but stay visible when users query AI assistants for relevant topics.

Whether that distinction holds in practice depends on crawler honesty about their declared purpose, which is worth being sceptical about. But the intent is sound, and it's the direction most sophisticated publishers are moving.

Does any of this actually send you visitors?

Here's the number that reframes the whole discussion. AI search engines collectively drove 16 times more referral traffic in 2026 compared to 2024. That growth is real. It's also coming off a very small base.

Among AI-sourced referrals, the dominant AI assistant accounts for roughly 62 to 75% of total AI referral traffic depending on the study and vertical. The next few competitors split the rest. The crawlers that do the most crawling — the training bots — send essentially zero referral traffic. The search-mode bots that do send referrals are actually some of the gentler crawlers in terms of request volume.

So there's an inversion worth sitting with: high crawl volume is not a proxy for high referral traffic. If you're evaluating AI crawlers purely by how often they hit your server, you're looking at the wrong metric.

What should you actually do with this?

Get visibility on what's hitting you. Standard analytics platforms don't distinguish between a training crawler and an AI search bot. You need raw server logs, or a proxy-layer tool that captures user-agent strings at the edge, to see the actual breakdown — which crawlers, which paths, which request patterns. If you don't have that visibility, you're guessing at a question that now has a measurable answer.

Update your robots.txt with clear intent. A blanket disallow-all-AI rule makes sense if you're purely a content business worried about data extraction. For most sites it's probably too blunt — you'd be blocking the referral traffic upside along with the training risk. The selective approach — blocking training crawlers while allowing search-mode bots — is worth a conversation, though you'll need to audit the specific user-agent tokens and keep the list current as new crawlers emerge.

Don't trust your referral attribution numbers. Most AI assistants strip the HTTP Referer header when a user follows a link from their interface, so those clicks often appear as direct traffic in your analytics. If you notice direct traffic spikes on pages that AI crawlers heavily index, or traffic surges 2 to 6 weeks after a major crawl event, that's probably AI-driven. Distinct landing paths or UTM-coded links in your structured data can help you get cleaner signal.

The volume of AI bot traffic isn't going down. The interesting management question — which bots you want, doing what, under what terms — is one most sites still haven't answered yet.

Sources

  1. What AI Crawlers Actually Want (August 2026 Update)
  2. We Analyzed robots.txt Across the Network: Publishers Now Block Training Bots and Allow Answering Bots (September 2026 Update)
  3. AI Crawler Volume Growth 2022-2026: GPTBot +305% YoY, AI Bots Now 22% of All Bot Traffic
  4. AI Search Stats 2026: Market Share, Referral, and Citation Data