Table of contents
- Quick summary
- One field, fifteen files, verified today
- Why one operator publishes four files, not one
- The field nobody documents
- The file that flipped
- The files that never sleep, and the ones that are
- Bug class two: the files also move, and most checkers do not follow
- What this costs in practice
- Verification does not stop at the IP list
- Why this stopped being academic: identity is now a billing decision
- What AIScan checks, and what it still gets wrong
- How to check today, in under a minute
- Your next step
Every guide on verifying an AI crawler tells you to check its request against the operator's published IP range file. Almost none of them tell you when that file was last regenerated. The answer is sitting in the file itself, one field, and it took one curl per operator to read it across all fifteen files still being published in September 2026.
Quick summary
The honest picture is not "these files are stale." It is that staleness is unevenly distributed and invisible from outside. Google regenerates four separate range files daily. OpenAI's chatgpt-user.json grew from 207 to 230 entries between 4 and 25 September 2026. And the file that would have been this article's headline example three weeks ago, OpenAI's gptbot.json, sat frozen for roughly eleven months and then regenerated on 22 September 2026, four days before this was fetched. A snapshot from three weeks ago is already wrong about which files are the stale ones.
| Regime | Example file | What it tells you |
|---|---|---|
| Regenerates daily | Google's four crawler-IP files | Fresh by definition; check the date anyway |
| Regenerates on its own schedule | OpenAI gptbot.json, Apple applebot.json | Long gaps are normal, not a sign of abandonment |
| Effectively frozen | Perplexity perplexitybot.json, Mistral mistralai-user-ips.json | Over nineteen months since the last change, as of today |
Every AI crawler IP file publishes a creationTime (or equivalent) field alongside the prefix list. It is the regeneration date of the file, not a promise about the ranges inside it, and reading it costs one request.
One field, fifteen files, verified today
The method is deliberately boring. Fetch each operator's published range file with curl -L (several of these URLs redirect, and a plain curl without -L silently reads a small stub instead of the file), parse the JSON, and print two numbers: how many prefixes it holds and the creationTime timestamp inside it. No login, no API key, about fifteen requests total.
Fetched from each operator's own domain and verified on 26 September 2026:
| File | Prefixes | Last regenerated | Age |
|---|---|---|---|
openai.com/gptbot.json | 18 | 22 Sep 2026 | 4 days |
openai.com/searchbot.json | 39 | 2 Jan 2026 | ~9 months |
openai.com/chatgpt-user.json | 230 | 25 Sep 2026 | 1 day |
openai.com/adsbot.json | 2 | 12 May 2026 | ~4.5 months |
claude.com/crawling/bots.json | 26 | 18 Aug 2026 | ~5.5 weeks |
perplexity.ai/perplexitybot.json | 8 | 7 Feb 2025 | over 19 months |
perplexity.ai/perplexity-user.json | 4 | 17 Oct 2025 | ~11 months |
Google common-crawlers.json | 317 | 25 Sep 2026 | 1 day |
Google special-crawlers.json | 272 | 25 Sep 2026 | 1 day |
Google user-triggered-fetchers.json | 1,058 | 25 Sep 2026 | 1 day |
Google user-triggered-agents.json (Google-Agent) | 20 | 25 Sep 2026 | 1 day |
search.developer.apple.com/applebot.json | 24 | 15 Sep 2026 | ~11 days |
index.commoncrawl.org/ccbot.json | 5* | 11 Aug 2026* | unverified today |
mistral.ai/mistralai-index-ips.json | 2 | 19 Apr 2026 | ~5 months |
mistral.ai/mistralai-user-ips.json | 4 | 19 Feb 2025 | over 19 months |
*Common Crawl's endpoint returned a connection failure (curl exit 35, HTTP 000) on three attempts today rather than a fresh read. That is a known intermittent fault on this specific host, not evidence of a stale file, so the last confirmed figures are carried forward and flagged rather than treated as current. www.anthropic.com/claudebot.json and mistral.ai/mistralai-training-ips.json were checked and remain hard 404s: neither exists to go stale.
Why one operator publishes four files, not one
Google alone accounts for four of the fifteen files, and the split is not arbitrary padding, it maps to a distinction Google's own crawler-verification documentation draws explicitly. According to that page, fetched on 26 September 2026, Google's crawlers and fetchers fall into three named categories, each with its own file: "common crawlers" (Googlebot and the other automated crawlers that "always respect robots.txt rules"), "special-case crawlers" (product-specific fetchers such as AdsBot that "may or may not respect robots.txt rules"), and "user-triggered fetchers" (tools like Site Verifier that act because a person asked them to, not on a schedule). The fourth file, user-triggered-agents.json, covers the newer Google-Agent identity for agentic browsing and is not part of that original three-way split, which is consistent with it being the file that grew fastest, from 4 prefixes to 20 in the same window the other three held their counts steady.
The practical upshot: "does Google publish a current IP file" is not one yes-or-no answer. A verifier built against common-crawlers.json alone will correctly validate Googlebot and miss every AdsBot or Site Verifier request, because those live in different files with different content and, this article's whole point, potentially different regeneration schedules. Checking one file and generalizing to "Google's ranges are current" checks a quarter of the actual surface.
The field nobody documents
creationTime (or, on a few files, a differently-named but equivalent timestamp) sits inside every one of these fifteen JSON files, and none of the operators' own crawler-verification pages mention it. Google's verifying-Googlebot documentation, fetched in full on 26 September 2026, explains what each file contains and how to match a prefix against it; it says nothing about reading the file's own regeneration date. OpenAI's public GPTBot documentation, checked the same day, is the same: it tells a publisher how to block or allow the crawler and does not mention that the range file it links to carries a timestamp at all.
That gap is why this measurement had to be done by hand rather than quoted from a guide: there is no operator-published guidance on how to interpret the field this whole article turns on. The convention exists because whoever built each file's generator happened to timestamp it, not because any operator considers it part of the public verification contract. Reading it is a courtesy the file's format allows, not a feature its documentation promises will always be there or always mean the same thing from one operator to the next.
The file that flipped
This table would have looked different on 7 September 2026, and the difference is the whole point. According to a research pass fetched from the same file on 7 September 2026, gptbot.json had not been regenerated since 30 October 2025, ten months earlier, and it was the single clearest example of an abandoned file: same 21 prefixes, same timestamp, checked three separate times over three weeks. As of today it carries 18 prefixes and a 22 September timestamp. Someone at OpenAI regenerated it four days before this article was fetched, after roughly eleven months of silence.
Nothing about that is visible from the file's structure. There is no changelog, no version number, no way to tell from the JSON itself whether a creationTime of nine months ago means "this hasn't needed to change" or "nobody is maintaining it" until it moves. A file that looked abandoned in September looks current in September of the same month, and the only way to know which state you are reading is to check it the moment before you rely on it.
The files that never sleep, and the ones that are
Set the flip aside and the fifteen files split into two populations that do not overlap. Google's four range files and OpenAI's chatgpt-user.json regenerate on something close to a daily cadence: the user-triggered-agents.json file (the Google-Agent identity) went from 4 prefixes in a March reading to 20 today, and chatgpt-user.json added 23 prefixes in three weeks. These are operators actively expanding infrastructure and republishing the evidence of it.
Then there is Perplexity's perplexitybot.json, unchanged since 7 February 2025, and Mistral's mistralai-user-ips.json, unchanged since 19 February 2025. Both are past nineteen months. Whatever changed about either company's crawling infrastructure since early 2025, none of it reached the file a publisher is told to check. Mistral's training IP file does not exist at all; it returns a 404, confirmed again today. A missing file and a nineteen-month-old file produce the identical outcome for anyone trying to verify a request: neither tells you anything current.
The lesson is not "trust Google, distrust Perplexity." It is that a single publishing convention, a JSON file of IP prefixes with a timestamp, is being used by nine or ten different organizations on completely different internal schedules, and nothing in the convention itself signals which schedule you are looking at.
Bug class two: the files also move, and most checkers do not follow
Staleness inside the file is one failure mode. The other is the file moving entirely. Google's historical URL, developers.google.com/search/apis/ipranges/googlebot.json, now 301s to developers.google.com/static/crawling/ipranges/common-crawlers.json, a different path and a different filename. All three of Google's crawler files relocated the same way. Perplexity's files moved from www.perplexity.com to www.perplexity.ai behind a 302.
Both old URLs still resolve, through the redirect, so this only breaks a checker that does not follow one. That describes most naive implementations: a script that reads the response body directly from a requests.get() or fetch() call without allow_redirects=True (or the language equivalent) gets whatever the old host serves at that path today, which in Google's case is a redirect stub of a few hundred bytes that parses as nothing useful, and in the worst case as an empty allowlist. A verifier that silently allowlists nobody fails closed and looks, from the outside, exactly like a correctly working check that happens to have found no valid ranges. Nobody gets an error. The IP just never matches, forever, and the mismatch is indistinguishable from a real spoofing attempt until someone thinks to check the raw response.
This is not a hypothetical. The research pass behind this article made the same mistake once: the first sweep of these fifteen files, run without -L, mis-recorded two operators as having moved to an empty or broken endpoint, when both were serving a normal file one redirect away. The fix was one flag. Finding it required noticing that two "broken" files belonged to operators with no history of publishing broken files, which is a slower way to catch a bug than just always following redirects in the first place.
What this costs in practice
Most sites do not hand-check an IP file before every request; they bake a snapshot into a WAF rule, an nginx allow block, or a scheduled job that regenerates an allowlist weekly or monthly. That gap between "when the file changes" and "when the allowlist regenerates" is where the two failure modes in this article turn into real outcomes rather than curiosities.
Against Google's four files, which regenerate close to daily, a monthly cron job is stale for most of the month by construction, and the fix is straightforward: shorten the interval or fetch on demand. Against Perplexity's perplexitybot.json, frozen for over nineteen months, a monthly cron job wastes a request checking a file that has not moved since early 2025, but it does no harm. The failure that actually costs something is the migration: a site that pinned developers.google.com/search/apis/ipranges/googlebot.json into a config file, rather than following the redirect at request time, is now reading a URL that still returns 200 but, depending on how the fetch was written, may be reading a stub rather than the live file it thinks it has. The allowlist looks unchanged in the config; the thing it actually resolves to has moved underneath it.
The asymmetry matters for prioritization: staleness inside a rarely-changing file is low-risk, and staleness caused by an unfollowed redirect is high-risk, because it can silently zero out an allowlist that a dashboard still reports as "loaded successfully." Any automated verification pipeline built against these files should log the resolved URL and the prefix count on every fetch, not just a success flag, specifically so a migration like Google's shows up as a visible change instead of a silent one.
Two distinct failure modes are worth telling apart, because the fix for each is different:
| Failure mode | What it looks like | How to catch it |
|---|---|---|
| Stale contents | Prefixes unchanged for months; creationTime far in the past | Read creationTime, not just the file's existence |
| Moved entirely | Old URL 301s or 302s to a new path; a non-following client reads an empty stub | Always fetch with -L; confirm the final URL, not just a 200 |
Verification does not stop at the IP list
An IP match is the first of two checks most guides recommend; the second is reverse DNS, and it does not cover every crawler either. Every documented AI crawler user-agent, and how to verify it carries the full thirty-three-token inventory and the reverse-DNS results for each one, and it is worth reading in full before building a verifier around either method alone, because the coverage gaps do not line up: some crawlers that publish a good IP file have no reverse-DNS hostname to check, and at least one range (a Google Cloud block also used by Anthropic-run infrastructure) passes a reverse-DNS check for the wrong operator entirely.
Six of ninety-one publisher hosts checked in an earlier pass, dated 7 September 2026, refused a self-declared Googlebot/2.1 request outright, five with a 403 and one with a 429, while serving an ordinary browser normally, meaning they verify by something other than trusting the header. That is the correct instinct. It is also evidence that this stopped being a theoretical exercise: real infrastructure is already checking, and a site that skips verification is depending on nobody testing it, which is a worse position than the alternative of checking and finding gaps.
None of this article's measurement extends to whether the prefixes inside a fresh file are themselves accurate, only to whether the file's own metadata says it was recently regenerated. A file can be freshly timestamped and still miss a range an operator brought online an hour earlier; regeneration cadence is a floor on trustworthiness, not a ceiling. Confirming prefix-level accuracy would require comparing a file's ranges against live traffic from that operator, which is a different, harder measurement this piece does not attempt.
Why this stopped being academic: identity is now a billing decision
According to Cloudflare's own Pay Per Crawl documentation, it bills per successful retrieval and settles the payment on a Web Bot Auth signature. An unsigned request offering the exact asking price gets 403 PaymentFailed, and the Discovery API refuses an unverified caller with 403 {"error":"Only verified bots can use this endpoint"}. Cloudflare Pay Per Crawl: should you charge AI crawlers? covers the mechanics in full; the relevant point here is narrower: the same identity question this article is measuring, which IP belongs to which crawler, and how current that mapping is, is no longer only a bot-access decision. It is the input to whether a request gets charged, refused, or served for free. A stale allowlist used to mean a missed crawler. It can now mean an operator paying for traffic your published ranges say does not exist, or your own site refusing a legitimate paying crawler because the file you checked it against was nineteen months old.
Stealth crawling vs. user-driven fetching measured a related edge on 75 real sites: a Disallow: / aimed at a training crawler is backed by an actual server-side refusal only 45 to 67 percent of the time, depending on the operator. Declaring a rule and enforcing it are two different projects, and the IP files this article measures are the enforcement half for the small number of publishers who do more than write a robots.txt line and hope. Whether a declared block actually stops training on already-published content is a separate question again, covered in Does blocking AI training protect your content?, but it rests on the same gap: a declaration is not a verified outcome.
What AIScan checks, and what it still gets wrong
AIScan's bot-access dimension covers both checks named in this article's tags. B2 reads whether a site's robots.txt names any known AI agent at all, which is a one-time, static read of a file the site controls. B3 checks for a published Web Bot Auth key directory, the mechanism the previous section just described as a live payment gate. According to AIScan's own production database, queried live across 739 real sites on 26 September 2026, B2 passes on 40.1 percent and B3 on 7.7 percent, both dimensions AIScan can grade purely from what a site publishes.
What AIScan cannot see is any of the fifteen files this article fetched by hand. Verifying that a specific inbound request actually originated from the IP range an operator currently publishes requires access to that request's source address at the moment it arrives, which is server-log or edge-log data no external scanner is ever handed. A 100/100 AIScan score is a statement about what a site declares and how it responds to a probe, not a guarantee that its edge is checking incoming crawler traffic against a current file. Those are different questions, and conflating them is exactly the mistake this whole article is about: trusting a static signal to answer a question that only a fresh check can answer.
B3 has a bug worth naming rather than hiding, because it is one more instance of the same problem: a piece of published metadata giving no way to tell current from stale. Grouped by evidence string across the same 739-site read, B3 reports the literal string HTTP 200 for 57 sites marked pass and the identical literal string HTTP 200 for 51 sites marked partial. The verdict differs; the evidence given for the verdict does not. First published against a 499-site read on 9 September, still reproducing today. Anyone who reads a check's status without reading what actually backs it, ours included, is trusting a label instead of checking the file, which is precisely the habit this article is arguing against.
How to check today, in under a minute
The fastest way to see where a specific site stands on this is to run it through AIScan: paste the URL at aiscan.site or run npx aiscan-cli yoursite.com, no account required, and read the B2 and B3 rows directly. That answers what the site declares and publishes.
To check an operator's IP file yourself, by hand, the command is one line per file:
curl -sL https://openai.com/gptbot.json | python3 -c "import json,sys; d=json.load(sys.stdin); print(d.get('creationTime'), len(d.get('prefixes', d.get('ips', []))))"
Swap in any of the fifteen URLs above. -L is not optional; three of them redirect. If a file's creationTime is months old, that is not automatically a problem, some operators genuinely do not need to change often, but it means the file is telling you what it looked like months ago, and you are the one deciding whether that is still good enough for what you are about to trust it for.
Because this article's own headline finding is that a three-week-old snapshot already misrepresented one of fifteen files, treat every number in the table above the same way: current as of 26 September 2026, not current forever. The fix is not to memorize which operators are fast and which are slow, it is to make the one-line check above part of whatever process relies on these files, on a schedule that matches how often the file in question has actually moved historically, monthly for the daily-regenerating ones is overkill, and a single annual check for a file frozen nineteen months is still a check worth doing, because nineteen months is exactly long enough for an operator to finally move it without anyone downstream noticing.
For a WordPress site making any of these declarations, ThinkRank manages robots.txt, robots meta tags, schema and llms.txt from one plugin rather than three separate ones fighting over the same file, and it migrates existing settings from Rank Math, Yoast, All in One SEO or SEOPress, so adopting it costs nothing already configured. For a Shopify store weighing the same AI-bot-access questions, StoreSEO covers the llms.txt and agents.md half of the same problem for commerce sites specifically.
Your next step
Run your own site through AIScan and read the B2 and B3 rows: aiscan.site or npx aiscan-cli yoursite.com. Then pick the one operator whose crawler matters most to your traffic and check its creationTime field directly, today, with the one-line command above. A file that was fresh in a previous article, including this one, is not a fact about the file, only about the day it was read. For everything else this dimension covers, the full checklist lives at aiscan.site/guides.
Frequently asked questions
My allowlist stopped matching GPTBot requests that used to pass. What changed?
OpenAI's gptbot.json was frozen at 21 prefixes from 30 October 2025 until it regenerated on 22 September 2026, now carrying 18. If your allowlist was pinned to a snapshot from before that date, a legitimate GPTBot request from a newer range will no longer match. Re-fetch the file with curl -L and rebuild the allowlist rather than assuming the mismatch means spoofing.
The Google crawler IP file I bookmarked now returns something that doesn't parse. What's wrong?
Google relocated all three of its classic range files from developers.google.com/search/apis/ipranges/ to developers.google.com/static/crawling/ipranges/, with new filenames as well as a new path. The old URL still returns 200 through a redirect, so a script that doesn't follow redirects reads a small stub instead of the real file and silently fails. Add -L to curl, or allow_redirects=True in code, and re-point at the new path.
My verification script shows zero IP matches for PerplexityBot even though it's clearly hitting my site. Is the file wrong?
Perplexity's perplexitybot.json has not been regenerated since 7 February 2025, over nineteen months. It may simply be missing a range Perplexity brought online since then. A zero-match result against a file this old is not proof of spoofing; it is a sign the file itself may be behind, and reverse DNS or a Web Bot Auth signature is a better second check for this specific operator.
How often should I re-fetch these IP range files?
Match the schedule to the file, not a fixed interval. Google's four files and OpenAI's chatgpt-user.json regenerate close to daily, so a weekly automated fetch is reasonable for those. Files that have gone a year or more without changing, like Perplexity's and Mistral's, still deserve a periodic check, monthly or quarterly is enough, specifically to catch the moment one of them finally moves.
Is a stale IP range file actually a security problem, or just an inconvenience?
Both, depending on which failure mode you hit. Stale contents inside an otherwise-working file mostly cause missed legitimate crawler traffic, which is an inconvenience. A silently broken redirect that zeroes out an allowlist is worse: it can pass a health check while blocking every request that relies on it, and now that Cloudflare's Pay Per Crawl bills on the same identity signal, a bad allowlist can also mean refusing a paying, legitimate crawler.
Does AIScan check these fifteen IP range files directly as part of a site scan?
No. AIScan's B2 and B3 checks read what a site itself publishes, its robots.txt declarations and any Web Bot Auth key directory, not the operator-side files this article measures. Verifying an inbound request's IP against a current operator file requires access to that request's source address at the edge, which is server-log data no external scanner can see.
My reverse-DNS check fails for GPTBot even though I'm confident the request is genuinely from OpenAI. Is my check broken?
Probably not. GPTBot, ChatGPT-User and ClaudeBot have no published reverse-DNS hostname to check against, so forward-confirmed reverse DNS has no method for them at all, pass or fail. That gap is documented in full in AIScan's crawler user-agent guide; for these specific operators, the IP range file (checked for freshness the way this article describes) is the more useful signal, not reverse DNS.
What's the fastest way to find out where my own site actually stands on AI bot verification?
Paste the URL at aiscan.site or run npx aiscan-cli yoursite.com and read the B2 and B3 rows; that takes under a minute and needs no account. To go one level deeper on a specific operator, run the one-line curl command in this article against that operator's file and read its own creationTime field.
Related guides
Audit Your Own Agent Readiness From a Terminal in Ten Minutes
Paste your URL at aiscan.sitehttps://aiscan.site/ or run npx aiscancli yoursite.com and you get a score in under thirty seconds. Fine for a first read, not enough if you actually want to know what…
The AI-Crawler Statistics Everyone Quotes, and Which You Can Actually Check
Every AIreadiness article, including plenty of ours, opens with a number: 78% of sites have a robots.txt, 28% publish an llms.txt, AI crawlers can't read JavaScript, one crawler is spoofed 17 times…
Web Bot Auth: Cryptographic Crawler Identity, Explained
Across the 499 sites in our own scan corpus, exactly 19 publish a working Web Bot Auth key directory. That's 3.8%. If you run any kind of automated client that fetches other people's sites a scraping…
Does Blocking AI Training Protect Your Content?
Too long, didn't read? Here's the honest version. | If you're asking... | The evidence says | What to do instead | |||| | Will blocking crawlers remove content already used for training? | No.…
