Key takeaways
- Cloudflare verifies crawler identity against operator-published IP ranges before counting a request, which is genuinely better than parsing raw logs full of spoofable user-agent strings. But full raw log export (Logpush) is Enterprise-only, with Business capped at 72-hour retention and Free/Pro getting aggregated analytics only.
- Independent measurement found 62.1% of requests claiming to be a known AI crawler couldn't be matched to the vendor's published IP range, and unconfirmed "AI crawler" traffic correlated heavily with scanning for
.envfiles and admin panels, not real bots. - Dedicated tools like Botify LogAnalyzer, DarkVisitors, and Screaming Frog's Log File Analyzer give you deeper historical analysis and page-level crawl coverage, but most are analytics-only: no blocking, no rate limiting, no Pay Per Crawl.
- The crawl-to-refer ratio is the metric that actually matters for deciding what to do with a crawler. ClaudeBot has been measured crawling over 11,000 pages per referral in some weeks, versus roughly 5:1 for Googlebot.
- Nobody's raw log data, including Cloudflare's own Radar numbers, will match what a GEO platform like Promptwatch reports, because the underlying datasets and verification methods differ. Treat them as complementary, not as a sanity check on each other.
Why this comparison keeps coming up
I get why people want a simple answer here. You're already paying for Cloudflare, it sits in front of your traffic, and it has a bot management product with a fancy ML fingerprinting engine. So why would you also pay for a dedicated AI crawler log tool?
The honest answer is that Cloudflare Bot Management and a dedicated AI crawler log tool are not really solving the same problem, even though they're both staring at the same firehose of GPTBot and ClaudeBot requests. Cloudflare's job is defense: allow, challenge, charge, or block. A dedicated log tool's job is usually analysis: which pages got crawled, how often, by whom, and what happened after.
On June 3, 2026, Cloudflare Radar reported that automated requests crossed 57.5% of all HTML web traffic. Bots now genuinely outnumber humans on the open web, according to Hydrolix's reporting on that milestone. That's the backdrop for this whole debate: you can't eyeball your logs anymore, you need a system, and the question is which system.
What Cloudflare actually gives you
Cloudflare's AI-specific product is called AI Crawl Control (renamed from AI Audit), and it's available on every plan including Free, with zero configuration required. The feature set is straightforward:
- See which AI services are accessing your content
- Set per-crawler allow/block policies
- Monitor robots.txt compliance and flag violators
- Pay Per Crawl, where you set a flat or dynamic per-request price and get paid when a crawler fetches a page
Pay Per Crawl is still labeled private beta in Cloudflare's docs as of mid-August 2026, but it's been live since July 2025, and Cloudflare added per-URI exemptions and dynamic pricing via a crawler-price header in June 2026. The company frames the whole thing as three levers: Allow, Charge, Block.

Cloudflare also publishes a public dashboard, Radar AI Insights, that breaks crawl purpose into training, search, and user-fetch categories, tracks crawl-to-refer ratios, and segments by industry. It's a genuinely useful public resource even if you're not a Cloudflare customer, because the scale behind it is enormous. Cloudflare sits in front of roughly 20% of internet traffic, so its bot fingerprinting dataset is hard for anyone else to match in raw size.
The catch: your own logs are gated by plan
Here's where the "but" comes in, and it's a big one. Full raw HTTP log access through Logpush is Enterprise-only. The Business plan gives you just 72 hours of raw log retention. Free and Pro customers get aggregated analytics, not the actual request-level data. One reviewer put it bluntly: they were shocked Cloudflare's business model means you can't fully download your own website's server logs unless you pay enterprise money, often cited around $5,000+/month for bot-management bundles on top of the base plan.
Compare that to Fastly, which offers real-time log streaming on all paid plans, not just its top tier. If raw log ownership matters to you, and for serious AI crawler analysis it usually does, Cloudflare's tiering is a real limitation, not a minor inconvenience.
What dedicated AI crawler log tools do differently
Dedicated tools generally fall into two camps: ones that sit as a layer in front of your site like a reverse proxy, and ones that ingest log files you already have. Either way, their pitch is depth: classify every request by LLM source, track crawl frequency and page-level coverage over time, and surface crawl-budget waste that generic analytics never catches (crawlers don't execute JavaScript, so GA4 never sees them at all).
A few concrete examples from the current landscape:
- DarkVisitors offers a free tier capped at 100,000 events a month; past that, new events simply stop showing up in your analytics rather than triggering an error. It combines a JS tag for live detection with log-upload for historical review and a pre-built taxonomy of AI agents.
- Botify LogAnalyzer ingests raw logs directly from cloud storage, which matters at enterprise scale, and segments by bot type and crawl-budget waste. It has no block or rate-limit controls at all, it's pure analytics.
- Scrape.do's Log Analyzer is built for one-time historical audits: upload a log file, get an AI-bot breakdown by user-agent in under five minutes, no ongoing infrastructure commitment.
- Screaming Frog's Log File Analyzer is free for up to 500 URLs, then a flat annual license. It's SEO-oriented rather than AI-citation-specific but a lot of technical SEOs already have it.
Notably, several comparison writeups explicitly leave Cloudflare and Fastly out of "best AI crawler analytics tools" rankings, arguing neither is purpose-built for the depth this category now requires, despite both having capable bot management. That's a fair framing: Cloudflare manages bots, dedicated tools analyze them.
| Tool | Tracking layer | Blocks/rate limits | Raw log access | Starting price |
|---|---|---|---|---|
| Cloudflare AI Crawl Control | Network edge, IP-verified | Yes (Allow/Charge/Block) | Enterprise-only for full export | Free (paid plans for deeper features) |
| DarkVisitors | JS tag + log upload | No | N/A (own dashboard) | Free up to 100k events/mo |
| Botify LogAnalyzer | Raw log ingestion | No | You bring your own logs | Enterprise-tier pricing |
| Screaming Frog Log File Analyzer | Local log upload | No | You bring your own logs | Free up to 500 URLs, ~£259/yr unlimited |
| Fastly bot management | Network edge | Yes | Real-time streaming, all paid plans | Custom |
The verification gap nobody talks about enough
This is the part that actually changes how you should read any crawler statistic you see online, including Cloudflare's own numbers versus a dedicated tool's numbers versus your own raw access logs.
HUMAN Security found that 5.7% of all traffic carrying a known AI crawler or scraper user-agent across 16 tracked crawlers was spoofed. ChatGPT-User specifically was spoofed 16.7% of the time. An independent small-site measurement over a two-week window in July and August 2026 found something even starker: of 9,452 requests claiming to be an AI crawler, 62.1% of checkable claims couldn't be matched to the operator's published IP ranges. The kicker is what that unconfirmed traffic was actually doing: requests for .env files, credentials, and admin panels made up just 0.5% of confirmed crawler traffic but 41% of unconfirmed traffic. A lot of "AI crawler" hits in a plain access log are just vulnerability scanners wearing a GPTBot costume.
Cloudflare's advantage here is structural, not just a bigger dataset. It confirms a bot's identity against the operator's published IP ranges before the request even gets counted, filtering spoofed traffic out upstream rather than labeling it and hoping someone downstream notices. Promptwatch's crawler analytics use the same IP-verification logic, which is one reason its AI crawler traffic data is worth trusting over a raw-log parse that only checks user-agent strings. If a dedicated tool you're evaluating can't tell you whether it's doing IP verification or just string matching, ask before you buy, because the gap between those two methods is proportionally worse on small sites where there's less real crawling to dilute the forged traffic.
The metric that actually matters: crawl-to-refer ratio
Raw crawl counts are close to useless on their own. What you want to know is whether a crawler sends anything back. Hydrolix's bot-management research frames this as the crawl-to-refer ratio, and it draws a sharp structural line: a crawler building an LLM training dataset harvests content and sends zero referrals, while an AI search agent crawls the same pages and drives measurable traffic back.

The numbers back this up and they're not subtle. ClaudeBot has been measured crawling over 11,000 pages per referral in some reporting windows. Cloudflare Radar's own public tracking showed a week in April 2026 where ClaudeBot ran at roughly 13,528:1, OpenAI's crawlers at about 1,252:1, and Googlebot at just 5:1. A separate June 2026 reading put Anthropic near 4,580:1, GPTBot near 848:1, and Perplexity around 186:1. The exact ratio moves week to week, but the structural imbalance doesn't: AI extractors out-crawl referral-generators by orders of magnitude, consistently.
Hydrolix's practical recommendation: calculate the crawl-to-refer ratio for your top 10 AI crawlers by request volume, and treat anything above roughly 1,000:1 with no improving referral trend as a candidate for rate limiting or Pay Per Crawl.
What's actually changing in the crawler mix right now
This is worth flagging because any tool you pick needs to keep up with it. Promptwatch's data shows OpenAI's share of verified AI crawler requests fell from 94.8% the week of June 8-14, 2026 to 79.8% the week of August 31-September 6, 2026. The bigger surprise is Meta: Meta-WebIndexer went from about 2.2% to 37.8% of all tracked AI crawler requests between mid-July and August 9, 2026, a roughly 17x jump in under a month, becoming the single heaviest crawler in Promptwatch's logs. That lines up with reports that Meta is building its own web index so Meta AI doesn't have to lean on Google. Independently, Claude's citation crawler grew more than 100x over four months, from about 30 visits a day in mid-December 2025 to several thousand a day by mid-April 2026.
If your bot management setup, Cloudflare or otherwise, is still tuned around 2025's OpenAI-dominant assumptions, you're probably mis-prioritizing right now. See the full breakdown in Promptwatch's Meta-WebIndexer report and the Claude citation crawler data.
Why your numbers never match anyone else's
If you've ever compared your own Cloudflare dashboard to a third-party crawler report and gotten a different percentage, that's expected, not a bug. Cloudflare confirms a bot's identity against the operator's published IP ranges before a request enters its dataset, so spoofed traffic is filtered out upstream. A dedicated log tool that only parses raw access logs for user-agent strings is measuring a different population entirely, one that includes scanners and scrapers pretending to be GPTBot.
The same logic applies when you're comparing Cloudflare Radar's public numbers to a GEO platform's citation data. Promptwatch explicitly integrates with Cloudflare, Vercel, and Fastly logs rather than replacing them, positioning itself as a layer that connects crawl activity to actual citations and traffic, not a competing raw-log source. If you're trying to understand why AI search traffic isn't showing up where you expect, tools in the GEO software directory at bestgeosoftware.com are built to connect that crawl data to the content and prompt side of the equation, which Cloudflare's bot management was never designed to do.

Where the enterprise detection gap actually sits
Hydrolix surveyed 300 enterprise leaders for its 2026 State of AI Bots report and found a confidence-reality gap worth sitting with: 79% said they were confident they could detect bot activity, but only 23% had a proactive, comprehensive bot management strategy. Leaders estimated AI bots at roughly 17% of their traffic, while Imperva's 2026 Bad Bot Report put total automated traffic above 53% of all web traffic, 40% of that malicious. That's a 36-point perception gap between what teams think is happening and what's actually happening.
Only 33% of respondents said their WAF or bot tool blocked more than half of AI bot traffic over the prior 12 months, and even then it's often unclear whether the blocked traffic was beneficial or malicious. The barriers cited were budget constraints (40%), fragmented systems and insufficient visibility (30%), and being overwhelmed by telemetry volume (27%). None of that gets fixed by switching vendors alone; it gets fixed by someone on the team actually owning the crawl-to-refer analysis on a schedule.
So which should you actually use
For most teams on Cloudflare already, start with AI Crawl Control. It's free, it's IP-verified, and the Allow/Charge/Block controls plus Radar's public crawl-to-refer data give you a real baseline without buying anything new. The problem shows up the moment you need raw log export for deeper historical analysis or custom tooling, since that's locked behind Enterprise pricing that commonly starts around $5,000 a month.
If you need that depth without the Enterprise price tag, a dedicated tool like Botify LogAnalyzer or DarkVisitors fills the gap, especially for log-file audits or page-level crawl coverage reporting that Cloudflare's dashboards don't break out. Just know you're buying analytics, not enforcement; you'll still need Cloudflare, Fastly, or a WAF to actually act on what you find.
And if your real question isn't "who's crawling me" but "is any of this crawling turning into AI search visibility and traffic," that's a different job entirely, one that crawler logs alone can't answer. That's where connecting crawler data to citation tracking, content gap analysis, and actual AI search traffic attribution matters, which is the layer Promptwatch sits at. For a broader look at how 21 platforms in this space stack up, the comparison at Promptwatch's GEO platform rundown is worth a read, and if you want to browse more options first, the AI rank tracking tools directory at ai-rank-tools.com covers a wider set of monitoring tools worth considering before you commit to any one stack.