Illustration of six rungs on a reproducibility ladder for AI-crawler statistics, from a green checkmark rung down to a coral not-reproducible rung, on a dark green background, aiscan.site
Illustration of six rungs on a reproducibility ladder for AI-crawler statistics, from a green checkmark rung down to a coral not-reproducible rung, on a dark green background, aiscan.site
AI Readiness

The AI-Crawler Statistics Everyone Quotes, and Which You Can Actually Check

We tested 10 widely-quoted AI-crawler statistics against their live sources on 27 Sep 2026. Some check out in one command; two had already vanished entirely.

AAsif Rahman 27 Sept 2026 18 min read
#AI crawlers#research methodology#AI SEO statistics#llms.txt

This guide covers B1 · Bot Access, B2 · Bot Access, C1 · Content, C2 · Content.

Table of contents

Every AI-readiness article, including plenty of ours, opens with a number: 78% of sites have a robots.txt, 28% publish an llms.txt, AI crawlers can't read JavaScript, one crawler is spoofed 17 times for every real request. The numbers get repeated for months. Almost nobody says whether a reader can check them.

We tried to check ten of the most-quoted ones ourselves, on 27 September 2026, using nothing but curl and a browser. Some took one command. One turned out to be gated behind the most expensive plan the vendor sells. Two we had already printed in our own back catalog and had to delete, because once the source became readable again, the number wasn't in it.

This matters more than a footnote about citation hygiene. A statistic that can't be reproduced is doing the same job as an opinion dressed up as a fact: it can't be checked, so it can't be argued with, and it keeps getting forwarded past the point where anyone remembers where it came from. A generative model asked to write about AI-crawler behavior will happily repeat "AI crawlers can't render JavaScript" or "70,900:1" as settled facts, because both circulate widely enough to look settled. Neither is, and the difference between the two is exactly the ladder below. We ran the same test on our own AI-crawler IP-range citations last week; see how stale those files actually are for the sibling result on files instead of blog claims.

Quick summary

StatisticSourceCan you check it yourself?What we found
Robots.txt on 78% of sites, Content Signals on 4%, Markdown negotiation on 3.9%Cloudflare Radar, published 17 Apr 2026Not against a live number; the post promises weekly updates and the page has not been edited since 15 Jul 2026Stale by 74 days on a chart that says it refreshes weekly
llms.txt on 28% of sites; 97% of those files got zero requestsAhrefs, ~137,000 domainsYes, with the sample-bias caveat Ahrefs states itselfThe 28% figure is an upper bound, by the authors' own words
Anthropic crawls 70,900 pages per referralCirculated widely, attributed to CloudflareNo; it does not appear in any Cloudflare post we can readDeleted from our own cache on 7 Sep 2026 after a direct search
AI crawlers can't render JavaScriptVercel + MERJ, 17 Dec 2024Only as a 2024 snapshot; no operator has confirmed or denied it sinceRoutinely quoted as if it were current
Average spoof ratio 1:17 for AI-crawler trafficHUMAN Security, 9 Sep 2025No; the source page returns HTTP 403 to every automated fetchFrozen at a 2025 reading with no way to refresh it
$24M in 30-day agent-payment volumex402.org homepage counterNo; it's a number baked into the page, not a live queryByte-identical across days we checked weeks apart

A reproducibility ladder, not a fact-check

A fact-check asks whether a number was ever true. That's a lower bar than this article is trying to clear. We're asking a narrower, more useful question: given nothing but the source URL, can a reader with a terminal get the same number today that a post is quoting?

That splits into six answers, and we sorted our ten candidates into them.

RungWhat it meansExample from this post
1. Reproducible in one commandA live page or public API returns the number todayCloudflare Radar's robots.txt and Content Signals figures
2. Reproducible with effort, source admits the biasA private dataset, but the sample and its skew are namedAhrefs' 28% llms.txt figure
3. Frozen, honestlyThe source page 403s every automated fetch todayHUMAN Security's 1:17 spoof ratio
4. Not reproducible at allNo traceable source, or the feature is gated behind a paid tierCrawl-to-referral ratios (Enterprise-only)
5. Dated wrong by everyoneReproducible in principle, but the date keeps getting stripped"AI crawlers can't render JavaScript" (Dec 2024)
6. Reproducible only if you re-probe itOne fetch shows an event; a policy needs several, spaced outThe matched-pair crawler-identity sweep
  1. Reproducible in one command. The source publishes a live page or a public API, and fetching it now gives you the number.
  2. Reproducible with effort, and the source says so. The number came from a private dataset, but the authors named their sample and its bias.
  3. Frozen, honestly. The source page is unreachable to an automated reader today, so the number can only ever be quoted at its original date.
  4. Not reproducible at all. The number circulates with no traceable source, or the feature that would produce it is gated behind a paid tier nobody publishing the number has access to.
  5. Dated wrong by everyone. The number is reproducible in principle, but every citation strips the date and presents a two-year-old snapshot as current.
  6. Reproducible only if you re-probe it. A single fetch tells you what happened once; the claim is really about a policy, and policies need a second and third fetch, spaced out, before you can tell a real pattern from a fluke.

Reproducible in one command, but the chart hasn't moved in ten weeks

Cloudflare's Agent Readiness post is the cleanest case of a genuinely live, first-party number: robots.txt on 78% of the 200,000 most-visited domains it sampled, Content Signals on 4%, Markdown content negotiation on 3.9%, and MCP Server Cards plus API Catalogs together on fewer than 15 sites. The post says outright that the chart "will be updated weekly."

Verified on 27 September 2026 by fetching the page fresh, its own structured data reads datePublished 17 April 2026 and dateModified 15 July 2026. That is one edit in five months, and the last one is now 74 days old on a page that promises a weekly refresh. The three headline percentages in the live HTML are word-for-word what we recorded in this queue's cache on 7 September, ten weeks ago. Whatever the real figure is today, the published one is not the reproducible answer to "what's happening right now" that the "updated weekly" line implies.

There's a second problem underneath the first one, and it's sharper than staleness. Our own documentation-sites study measured 52 real documentation hosts against every way a client can actually ask for a Markdown twin: an Accept: text/markdown header, a .md suffix, and an /index.md path. The header alone found 29 of 34 Markdown-serving sites; the .md suffix found 27, five of which the header never reaches. Radar's 3.9% figure is a single-route measurement: it almost certainly probes the header only, the way our own C1 check did until we measured the gap. A single-route number for a multi-route convention isn't just old, it's the wrong shape of question.

Reproducible with effort, and the authors say so

According to Ahrefs' study of 137,000 domains, this is the honest version of a private-dataset claim. 28% of the sample publishes a valid llms.txt, and 97% of those files received zero requests in Ahrefs' own May-2026 web analytics. Fetched again on 27 September, the post still carries its own caveat, verbatim: "Ahrefs Web Analytics customers skew more technical and SEO-aware than the web at large, so treat the 28% adoption figure as an upper bound." That single sentence is doing more honest work than most posts that quote the 28% number do, because almost none of them repeat the caveat.

Rankability's tracker, refreshed monthly against a named Tranco list, puts the top-1,000 sites at 8.7% (15.8% of the ones it could actually reach) and the top-10,000 at 5.6%. SE Ranking's ~300,000-domain sample lands at 10.13% and adds a finding nobody else has: llms.txt added no predictive value to its citation model, and removing it as a feature improved the model's accuracy.

Three samples, three different denominators, three different numbers (5.6%, 10.13%, 28%), and every single one of them is honestly a fact about llms.txt on a particular population, not a fact about the web. Our own corpus adds a fourth: adoption on the friendly, technical population of documentation sites specifically is 71% (37 of 52), the highest anyone has published, which is the ceiling all three general-web numbers are really being compared against without saying so. Our does-llms-txt-actually-work post has the fuller answer on whether any of this adoption translates into anything an AI assistant actually reads.

The SE Ranking figure carries a second finding worth more attention than the adoption number itself: llms.txt added no predictive value to its own AI-citation model, and removing the feature actually improved the model's accuracy. That's a company with a commercial reason to find llms.txt useful publishing a result that says it isn't, on their own dataset, and it hasn't been contradicted by a larger study since. A statistic that goes against the publisher's own interest is worth more weight than one that flatters it, and this is the one number in this whole survey with that property.

Not reproducible at all: the number that got deleted

The sharpest test of this whole exercise is a negative result, and it's ours. Two figures used to sit in our own research cache: "more than 50% of crawl traffic from good bots goes to re-fetching pages that haven't changed," and crawl-to-referral ratios of roughly Anthropic 70,900:1 and Mistral 0.1:1, both attributed to Cloudflare. On 7 September 2026, once blog.cloudflare.com article bodies became readable to a plain curl fetch again, we searched six candidate Cloudflare posts for both phrases and both numbers. Neither string, and neither figure, appears in any of them. We deleted both from our own cache rather than keep citing something we could no longer locate.

The underlying concept is real and first-party: Cloudflare documents Attribution Business Insights, "a dashboard showing site-wide and per-operator crawl-to-referral ratios alongside bot traffic to your content," shipped 1 July 2026. It's gated to Enterprise Bot Management customers only. So the specific ratios that circulated for months were either an unpublished internal figure someone quoted from a briefing, or a hallucination that outlived its own source. Either way, nobody outside an Enterprise Cloudflare account can check a crawl-to-referral ratio for any site today, including their own. That gate is itself the publishable fact. A number nobody can reproduce, sitting behind a paywall nobody who's quoting it has access to, is a different kind of unreliable than a stale chart.

Dated wrong by everyone

"AI crawlers can't render JavaScript" is the single most repeated claim in this field, and its source has a date that keeps getting dropped. Vercel and MERJ's study, "The rise of the AI crawler," was published 17 December 2024 and carries no update notice on the live page today. Verbatim: "The results consistently show that none of the major AI crawlers currently render JavaScript," covering OpenAI, Anthropic, Meta, ByteDance and Perplexity's crawlers. The paper's own caveat rarely survives the retelling: "content included in the initial HTML response... may still be indexed since AI models can interpret non-HTML content." Microsoft Copilot was excluded from the study entirely, because it lacked a trackable user agent at the time, so the honest phrasing is "as last measured in December 2024, and no operator has said otherwise since," never "AI crawlers cannot render JavaScript" as a settled, current fact.

A second, older Vercel/MERJ measurement gets confused with it: a July 2024 Googlebot render-rate study, over 37,000 pages, found 100% of HTML pages were fully rendered. That study is about Google's own crawler and its Gemini token, which inherits Googlebot's renderer. It has nothing to say about GPTBot, ClaudeBot or PerplexityBot, and conflating the two dates and the two crawlers into one "AI can/can't render JS" sentence is how a nearly two-year-old finding keeps sounding fresh.

Frozen, honestly: one page, one year, no follow-up

HUMAN Security's Satori Threat Intelligence report, published 9 September 2025, measured a two-week window and found an average spoof ratio of 1:17 for traffic claiming to be an AI crawler, with spoofed requests at 5.7% of all AI-crawler-labelled traffic. Per user agent: ChatGPT-User 1:5, MistralAI-User 1:37, Perplexity-User 1:88.

Verified on 27 September 2026 with a plain automated request, the page still returns HTTP 403, the same result recorded when this figure first entered our cache. No 2026 equivalent exists that we can find. This is the honest version of "frozen": we're not saying the number is wrong, only that nobody, us included, can currently pull a fresh reading, so every citation of 1:17 is really a citation of a September 2025 snapshot that has had thirteen months to drift and cannot be checked against anything newer.

Reproducible only from a residential address

A whole category of "who blocks AI crawlers" statistics quietly depends on where the measuring computer sits, and almost nobody states their own refusal rate. Fetching from a datacentre IP, our own sweeps have returned wildly different exclusion rates depending on the population: 31% of 110 ordinary publisher homepages (7 September), 41% of 75 mixed hosts (9 September), but only 3.7% of 52 documentation services (10 September, our documentation-sites-ai-agent-readiness post). The refusal rate is a property of the population you sample, not of the measuring container: a hostile publisher list and a friendly docs-site list from the same container produce numbers ten times apart. Almost no published "N% of sites block AI bots" statistic states which kind of population it drew from, or the IP class it was measured with, and both change the answer more than the underlying policy does.

None of these three refusal rates is more correct than the others; they're answers to three different questions wearing the same headline shape. A study that says "34% of sites block AI crawlers" without naming its population is not comparable to one that says "3.7%", and treating them as competing measurements of one underlying truth is how a reader ends up more confused after reading two sources than after reading none.

A headline counter that never moves

x402.org's homepage displays a live-looking dashboard: trailing-30-day totals for transactions, volume, buyers and sellers. Fetched from the live page on 27 September 2026, it showed 75.41M transactions, $24.24M volume, 94.06K buyers and 22K sellers. Fetched twice in a row, seconds apart, the response body was byte-for-byte identical, and this queue first caught the same pattern comparing readings taken two days apart, in early September. The numbers aren't computed on request; they're baked into the served HTML as a snapshot, refreshed on the site's own schedule rather than live. It reads exactly like a real-time counter and isn't one, which makes it the cleanest small example in this whole survey of the gap between what a number looks like and what a reader can actually verify by asking twice.

Reproducible only if you re-probe it

The trickiest category isn't about whether a source is honest. It's about whether a single fetch can even answer the question being asked. Several of the most useful numbers this queue has produced are policies, not facts, and a policy only shows itself across repeated, spaced observations.

Our own matched-pair crawler sweep sent an operator's training-crawler token and its user-driven-fetcher token to 75 hosts and found 16 hosts that answered the two identities differently. Every one of those 16 differentials was re-probed twice more, spaced apart, and every one reproduced 3 of 3, which is what turned "we saw a difference once" into "this is a deterministic policy" rather than a rate limit, an A/B test, or a CDN hiccup that happened to line up. The 52-documentation-site Markdown-route sweep did the same thing to its own method: 78 of 80 probe slots held identical across three spaced rounds, and the two that didn't turned out to be a missing curl -L in the first pass, not a change at the target. "Was this re-probed?" is a fair question to ask of any AI-crawler statistic making a policy claim, including several of ours before 9 September 2026, and by that standard, most published numbers in this field, ours included, are still ungraded on it.

The one check that catches most of this

Nine of the ten statistics in this article resolve with a single request, and it's the same request every time: fetch the source page's own metadata and read when it last changed.

curl -sI "https://blog.cloudflare.com/agent-readiness/" | grep -i last-modified
# or, for a page that publishes structured data instead of a header:
curl -s "https://blog.cloudflare.com/agent-readiness/" | grep -o '"dateModified":"[^"]*"'

A last-modified header or a dateModified field that predates the claim you're about to make is the single most common failure mode in this entire survey, and it took one request to catch on Cloudflare's own chart. Where a page offers neither, as x402.org does, the fallback is the two-fetch test: request the same URL twice, minutes apart, and diff the bodies. Identical bytes mean a snapshot, not a live figure, whatever the page's design implies.

What we're grading ourselves on, too

Self-Publisher Protocol applies here in full: we don't get graded on a curve just because we run the scanner. Our own scan corpus, first published in full in state of AI agent readiness, reads as of 27 September 2026 658 sites on rubric 2026.08.2, mean score 55.4, down slightly from the 56.4-56.5 range recorded across 473–501 sites through mid-September. That's not a trend; it's a rolling roughly-ten-day window rather than a fixed cohort, and every repeat of this population must say so rather than imply progress or decline.

Our own C1 check (Markdown content negotiation) is a live case of a number that needs re-derivation, not reprinting, every time. As of today it passes on 197 of 658 hosts (29.9%). Of those 197 passes, 82 (41.6%) carry an evidence string with no mention of text/markdown at all, almost identical to the 41.4% figure our documentation-sites-ai-agent-readiness post measured on a 501-host corpus 17 days ago, on a corpus that has since grown by a third. Two independent readings, 17 days and 157 hosts apart, landing 0.2 points apart is the kind of convergence that makes a number worth trusting; a single reading of either would not have been.

And one number we generate ourselves is honestly not reproducible at all: ai_referral_visits, our own analytics table for inbound clicks from an AI answer, holds zero rows, unchanged since it was first checked in early September. We do not know whether that is a real zero or an instrumentation gap in our own pipeline, and we're saying so here rather than pretending it's settled.

The standard we're holding everyone else to

Every first-party sweep this queue has run publishes its population, its date and its method, specifically so a stranger can rerun it:

SweepPopulationMethodDate
Conditional-request round trip (full writeup)60 public hosts, 4 URLs eachFetch twice, check for a 3045 Sep 2026
Identity-varied pricing sweep119 publisher hostsFetch as a named AI crawler UA vs. browser6 Sep 2026
Publisher robots.txt configuration96–110 readable filesParse RFC 9309 groups, no live probe7 Sep 2026
Storefront protocol sweep30 Shopify storefrontsProbe .well-known paths and POST /api/mcp8 Sep 2026
Matched-pair identity sweep75 hosts x 7 identitiesTraining token vs. user-driven token, re-probed 3x9 Sep 2026
Documentation Markdown-route sweep52 documentation services, 385 probes4 request shapes, re-probed 3 rounds10 Sep 2026
Full rubric corpus study658 hosts (rolling), one rubric versionSQL aggregation, dedupe to latest scan per hostOngoing, this post

None of these needs anything beyond curl, a spreadsheet and the population definition printed above each one. That is the bar every AI-crawler statistic should clear before a second blog post repeats it, and it's a bar our own back catalog didn't always clear either, which is why two numbers are gone from it.

What AIScan checks for you, and what it still can't

AIScan's B1 and B2 checks read Content Signals declarations and the AI-crawler Disallow/Allow rules in your own robots.txt, and C2 checks for a valid llms.txt at the scanned path (build one with the llms.txt generator if you don't have one yet). Running a scan at aiscan.site or npx aiscan-cli yoursite.com gives you your own current reading against all three in under a minute, which is a faster and more current answer than quoting anyone's April survey.

What none of those checks can tell you, and what this whole article has been about: whether a statistic about the wider web is still current, whether it's being asked in a way that actually reaches the phenomenon (Radar's single-route Markdown figure), or whether the source page has quietly stopped updating. A scanner grades a URL. It can't grade a claim in somebody else's blog post, and neither can any other tool in this category. That's a reading job, and this article is the checklist for doing it.

There's a reason to care about this beyond tidiness, and it's the same reason this queue exists at all: if an AI answer engine is drawing on the same corpus of repeated, undated statistics that a human reader would find with a search, then getting a number wrong doesn't just mislead one reader, it seeds the next summary that cites your page as a source. A page that dates its claims and states its sample is more useful to a machine reader for exactly the same reason it's more useful to a human one, and that is the whole case for doing the two-request check above before publishing a number rather than after someone points out it's gone stale.

For the WordPress half of what these numbers describe — robots.txt, robots meta, llms.txt and schema all currently need three or four separate plugins fighting over the same file: ThinkRank handles all four from one plugin and migrates cleanly from Rank Math, Yoast, All in One SEO or SEOPress. For the Shopify half, where none of B1, B2 or C2 is built into the platform at all, StoreSEO generates llms.txt and manages the AEO/GEO settings these statistics are actually measuring; it's rated 5.0 from 756 reviews on the Shopify App Store, checked the day this was published.

Run your own site through AIScan to see current B1, B2 and C2 results rather than relying on anyone's dated survey, and see the full method behind every sweep in this post on our guides page.

Frequently asked questions

Is Cloudflare's 78% robots.txt figure still accurate right now?

We can't tell you, and neither can Cloudflare's own page. It was measured in April 2026 against 200,000 domains and last edited on 15 July 2026, despite the post's promise of a weekly refresh. Treat any citation of 78%, 4% or 3.9% from that post as an April 2026 snapshot, not a live number.

Can I check a crawl-to-referral ratio for my own site?

Not without an Enterprise Bot Management plan on Cloudflare. The feature that produces that number, Attribution Business Insights, shipped 1 July 2026 and is gated to Cloudflare's most expensive tier. The specific ratios that circulated earlier in 2026 don't appear in any Cloudflare post we can read, and we deleted them from our own cache for that reason.

Why doesn't Ahrefs' 28% llms.txt adoption figure match Rankability's 5.6% or SE Ranking's 10.13%?

Because they're three different samples measuring three different populations, not three attempts at the same number. Ahrefs sampled 137,000 of its own Web Analytics customers and says outright that this group "skews more technical and SEO-aware than the web at large." Rankability uses a named Tranco top-10,000 list. SE Ranking used roughly 300,000 domains. None of the three is wrong; none of them is "the" adoption rate either.

Is it still true that AI crawlers can't render JavaScript?

As last measured, yes, but that measurement is from 17 December 2024, and no crawler operator has published anything newer to confirm or contradict it. If you see the claim stated as a current fact rather than a 2024 finding, that's the tell that the date got stripped somewhere along the citation chain.

Why can't I re-check the HUMAN Security spoof-ratio figures myself?

The source page returns HTTP 403 to an automated request, including ours, as of 27 September 2026. The 1:17 average spoof ratio and the per-crawler splits are frozen at a 9 September 2025 reading with no way to pull a fresh one from that source.

Is x402.org's "$24M in 30 days" a live number?

No. Fetching the page twice in immediate succession returns byte-identical HTML, which means the figures are a snapshot baked into the page rather than computed per request. Treat any number from that dashboard as dated to whenever you fetched it, not as continuously updating.

How do I actually verify a statistic like these myself?

Fetch the primary source directly, not a blog post quoting it, and check three things: does the page state a date or a "last updated" field, does the methodology name a sample size and a bias, and does a repeat fetch return the same live number or a cached one. If any of the three is missing, the statistic belongs in the "not currently reproducible" bucket, whatever the number itself says.

What's the fastest single check for whether an AI-crawler statistic can still be trusted?

Refetch the primary source and read its own dateModified or "last updated" field against today's date, the same check that caught Cloudflare's Radar chart at 74 days stale on a weekly-refresh promise. It's one request and it catches the single most common failure mode in this whole survey: a number that was accurate once and has simply stopped being re-verified by the people who published it.

Related guides