---
title: "How Stale Are AI Crawler IP Lists? We Checked All 15 Files"
slug: ai-crawler-ip-list-staleness
published: 2026-09-26T03:26:13.998393+00:00
updated: 2026-09-26T03:26:13.998393+00:00
author: "Asif Rahman"
author_url: https://masifrahman.com
category: "AI Readiness"
tags: AI crawlers, IP verification, robots.txt, Web Bot Auth, check:B2, check:B3
description: "We fetched all 15 published AI crawler IP range files and read each one's own regeneration timestamp. Some update daily; one was frozen for 19 months straight."
url: https://aiscan.site/blog/ai-crawler-ip-list-staleness
---

Every guide on verifying an AI crawler tells you to check its request against the operator's published IP range file. Almost none of them tell you when that file was last regenerated. The answer is sitting in the file itself, one field, and it took one `curl` per operator to read it across all fifteen files still being published in September 2026.

## Quick summary

The honest picture is not "these files are stale." It is that staleness is unevenly distributed and invisible from outside. Google regenerates four separate range files daily. OpenAI's `chatgpt-user.json` grew from 207 to 230 entries between 4 and 25 September 2026. And the file that would have been this article's headline example three weeks ago, OpenAI's `gptbot.json`, sat frozen for roughly eleven months and then regenerated on 22 September 2026, four days before this was fetched. A snapshot from three weeks ago is already wrong about which files are the stale ones.

| Regime | Example file | What it tells you |
|---|---|---|
| Regenerates daily | Google's four crawler-IP files | Fresh by definition; check the date anyway |
| Regenerates on its own schedule | OpenAI `gptbot.json`, Apple `applebot.json` | Long gaps are normal, not a sign of abandonment |
| Effectively frozen | Perplexity `perplexitybot.json`, Mistral `mistralai-user-ips.json` | Over nineteen months since the last change, as of today |

Every AI crawler IP file publishes a `creationTime` (or equivalent) field alongside the prefix list. It is the regeneration date of the *file*, not a promise about the ranges inside it, and reading it costs one request.

## One field, fifteen files, verified today

The method is deliberately boring. Fetch each operator's published range file with `curl -L` (several of these URLs redirect, and a plain `curl` without `-L` silently reads a small stub instead of the file), parse the JSON, and print two numbers: how many prefixes it holds and the `creationTime` timestamp inside it. No login, no API key, about fifteen requests total.

Fetched from each operator's own domain and verified on 26 September 2026:

| File | Prefixes | Last regenerated | Age |
|---|---|---|---|
| `openai.com/gptbot.json` | 18 | 22 Sep 2026 | 4 days |
| `openai.com/searchbot.json` | 39 | 2 Jan 2026 | ~9 months |
| `openai.com/chatgpt-user.json` | 230 | 25 Sep 2026 | 1 day |
| `openai.com/adsbot.json` | 2 | 12 May 2026 | ~4.5 months |
| `claude.com/crawling/bots.json` | 26 | 18 Aug 2026 | ~5.5 weeks |
| `perplexity.ai/perplexitybot.json` | 8 | 7 Feb 2025 | over 19 months |
| `perplexity.ai/perplexity-user.json` | 4 | 17 Oct 2025 | ~11 months |
| Google `common-crawlers.json` | 317 | 25 Sep 2026 | 1 day |
| Google `special-crawlers.json` | 272 | 25 Sep 2026 | 1 day |
| Google `user-triggered-fetchers.json` | 1,058 | 25 Sep 2026 | 1 day |
| Google `user-triggered-agents.json` (Google-Agent) | 20 | 25 Sep 2026 | 1 day |
| `search.developer.apple.com/applebot.json` | 24 | 15 Sep 2026 | ~11 days |
| `index.commoncrawl.org/ccbot.json` | 5* | 11 Aug 2026* | unverified today |
| `mistral.ai/mistralai-index-ips.json` | 2 | 19 Apr 2026 | ~5 months |
| `mistral.ai/mistralai-user-ips.json` | 4 | 19 Feb 2025 | over 19 months |

\*Common Crawl's endpoint returned a connection failure (`curl` exit 35, HTTP 000) on three attempts today rather than a fresh read. That is a known intermittent fault on this specific host, not evidence of a stale file, so the last confirmed figures are carried forward and flagged rather than treated as current. `www.anthropic.com/claudebot.json` and `mistral.ai/mistralai-training-ips.json` were checked and remain hard 404s: neither exists to go stale.

## Why one operator publishes four files, not one

Google alone accounts for four of the fifteen files, and the split is not arbitrary padding, it maps to a distinction Google's own crawler-verification documentation draws explicitly. According to that page, fetched on 26 September 2026, Google's crawlers and fetchers fall into three named categories, each with its own file: "common crawlers" (Googlebot and the other automated crawlers that "always respect robots.txt rules"), "special-case crawlers" (product-specific fetchers such as AdsBot that "may or may not respect robots.txt rules"), and "user-triggered fetchers" (tools like Site Verifier that act because a person asked them to, not on a schedule). The fourth file, `user-triggered-agents.json`, covers the newer `Google-Agent` identity for agentic browsing and is not part of that original three-way split, which is consistent with it being the file that grew fastest, from 4 prefixes to 20 in the same window the other three held their counts steady.

The practical upshot: "does Google publish a current IP file" is not one yes-or-no answer. A verifier built against `common-crawlers.json` alone will correctly validate Googlebot and miss every AdsBot or Site Verifier request, because those live in different files with different content and, this article's whole point, potentially different regeneration schedules. Checking one file and generalizing to "Google's ranges are current" checks a quarter of the actual surface.

## The field nobody documents

`creationTime` (or, on a few files, a differently-named but equivalent timestamp) sits inside every one of these fifteen JSON files, and none of the operators' own crawler-verification pages mention it. Google's verifying-Googlebot documentation, fetched in full on 26 September 2026, explains what each file contains and how to match a prefix against it; it says nothing about reading the file's own regeneration date. OpenAI's public GPTBot documentation, checked the same day, is the same: it tells a publisher how to block or allow the crawler and does not mention that the range file it links to carries a timestamp at all.

That gap is why this measurement had to be done by hand rather than quoted from a guide: there is no operator-published guidance on how to interpret the field this whole article turns on. The convention exists because whoever built each file's generator happened to timestamp it, not because any operator considers it part of the public verification contract. Reading it is a courtesy the file's format allows, not a feature its documentation promises will always be there or always mean the same thing from one operator to the next.

## The file that flipped

This table would have looked different on 7 September 2026, and the difference is the whole point. According to a research pass fetched from the same file on 7 September 2026, `gptbot.json` had not been regenerated since 30 October 2025, ten months earlier, and it was the single clearest example of an abandoned file: same 21 prefixes, same timestamp, checked three separate times over three weeks. As of today it carries 18 prefixes and a 22 September timestamp. Someone at OpenAI regenerated it four days before this article was fetched, after roughly eleven months of silence.

Nothing about that is visible from the file's structure. There is no changelog, no version number, no way to tell from the JSON itself whether a `creationTime` of nine months ago means "this hasn't needed to change" or "nobody is maintaining it" until it moves. A file that looked abandoned in September looks current in September of the same month, and the only way to know which state you are reading is to check it the moment before you rely on it.

## The files that never sleep, and the ones that are

Set the flip aside and the fifteen files split into two populations that do not overlap. Google's four range files and OpenAI's `chatgpt-user.json` regenerate on something close to a daily cadence: the `user-triggered-agents.json` file (the `Google-Agent` identity) went from 4 prefixes in a March reading to 20 today, and `chatgpt-user.json` added 23 prefixes in three weeks. These are operators actively expanding infrastructure and republishing the evidence of it.

Then there is Perplexity's `perplexitybot.json`, unchanged since 7 February 2025, and Mistral's `mistralai-user-ips.json`, unchanged since 19 February 2025. Both are past nineteen months. Whatever changed about either company's crawling infrastructure since early 2025, none of it reached the file a publisher is told to check. Mistral's *training* IP file does not exist at all; it returns a 404, confirmed again today. A missing file and a nineteen-month-old file produce the identical outcome for anyone trying to verify a request: neither tells you anything current.

The lesson is not "trust Google, distrust Perplexity." It is that a single publishing convention, a JSON file of IP prefixes with a timestamp, is being used by nine or ten different organizations on completely different internal schedules, and nothing in the convention itself signals which schedule you are looking at.

## Bug class two: the files also move, and most checkers do not follow

Staleness inside the file is one failure mode. The other is the file moving entirely. Google's historical URL, `developers.google.com/search/apis/ipranges/googlebot.json`, now 301s to `developers.google.com/static/crawling/ipranges/common-crawlers.json`, a different path and a different filename. All three of Google's crawler files relocated the same way. Perplexity's files moved from `www.perplexity.com` to `www.perplexity.ai` behind a 302.

Both old URLs still resolve, through the redirect, so this only breaks a checker that does not follow one. That describes most naive implementations: a script that reads the response body directly from a `requests.get()` or `fetch()` call without `allow_redirects=True` (or the language equivalent) gets whatever the old host serves at that path today, which in Google's case is a redirect stub of a few hundred bytes that parses as nothing useful, and in the worst case as an empty allowlist. A verifier that silently allowlists nobody fails closed and looks, from the outside, exactly like a correctly working check that happens to have found no valid ranges. Nobody gets an error. The IP just never matches, forever, and the mismatch is indistinguishable from a real spoofing attempt until someone thinks to check the raw response.

This is not a hypothetical. The research pass behind this article made the same mistake once: the first sweep of these fifteen files, run without `-L`, mis-recorded two operators as having moved to an empty or broken endpoint, when both were serving a normal file one redirect away. The fix was one flag. Finding it required noticing that two "broken" files belonged to operators with no history of publishing broken files, which is a slower way to catch a bug than just always following redirects in the first place.

## What this costs in practice

Most sites do not hand-check an IP file before every request; they bake a snapshot into a WAF rule, an nginx `allow` block, or a scheduled job that regenerates an allowlist weekly or monthly. That gap between "when the file changes" and "when the allowlist regenerates" is where the two failure modes in this article turn into real outcomes rather than curiosities.

Against Google's four files, which regenerate close to daily, a monthly cron job is stale for most of the month by construction, and the fix is straightforward: shorten the interval or fetch on demand. Against Perplexity's `perplexitybot.json`, frozen for over nineteen months, a monthly cron job wastes a request checking a file that has not moved since early 2025, but it does no harm. The failure that actually costs something is the migration: a site that pinned `developers.google.com/search/apis/ipranges/googlebot.json` into a config file, rather than following the redirect at request time, is now reading a URL that still returns 200 but, depending on how the fetch was written, may be reading a stub rather than the live file it thinks it has. The allowlist looks unchanged in the config; the thing it actually resolves to has moved underneath it.

The asymmetry matters for prioritization: staleness inside a rarely-changing file is low-risk, and staleness caused by an unfollowed redirect is high-risk, because it can silently zero out an allowlist that a dashboard still reports as "loaded successfully." Any automated verification pipeline built against these files should log the resolved URL and the prefix count on every fetch, not just a success flag, specifically so a migration like Google's shows up as a visible change instead of a silent one.

Two distinct failure modes are worth telling apart, because the fix for each is different:

| Failure mode | What it looks like | How to catch it |
|---|---|---|
| Stale contents | Prefixes unchanged for months; `creationTime` far in the past | Read `creationTime`, not just the file's existence |
| Moved entirely | Old URL 301s or 302s to a new path; a non-following client reads an empty stub | Always fetch with `-L`; confirm the final URL, not just a 200 |

## Verification does not stop at the IP list

An IP match is the first of two checks most guides recommend; the second is reverse DNS, and it does not cover every crawler either. [Every documented AI crawler user-agent, and how to verify it](https://aiscan.site/blog/ai-crawler-user-agent-list-2026) carries the full thirty-three-token inventory and the reverse-DNS results for each one, and it is worth reading in full before building a verifier around either method alone, because the coverage gaps do not line up: some crawlers that publish a good IP file have no reverse-DNS hostname to check, and at least one range (a Google Cloud block also used by Anthropic-run infrastructure) passes a reverse-DNS check for the wrong operator entirely.

Six of ninety-one publisher hosts checked in an earlier pass, dated 7 September 2026, refused a self-declared `Googlebot/2.1` request outright, five with a 403 and one with a 429, while serving an ordinary browser normally, meaning they verify by something other than trusting the header. That is the correct instinct. It is also evidence that this stopped being a theoretical exercise: real infrastructure is already checking, and a site that skips verification is depending on nobody testing it, which is a worse position than the alternative of checking and finding gaps.

None of this article's measurement extends to whether the prefixes inside a fresh file are themselves accurate, only to whether the file's own metadata says it was recently regenerated. A file can be freshly timestamped and still miss a range an operator brought online an hour earlier; regeneration cadence is a floor on trustworthiness, not a ceiling. Confirming prefix-level accuracy would require comparing a file's ranges against live traffic from that operator, which is a different, harder measurement this piece does not attempt.

## Why this stopped being academic: identity is now a billing decision

According to Cloudflare's own Pay Per Crawl documentation, it bills per successful retrieval and settles the payment on a Web Bot Auth signature. An unsigned request offering the exact asking price gets `403 PaymentFailed`, and the Discovery API refuses an unverified caller with `403 {"error":"Only verified bots can use this endpoint"}`. [Cloudflare Pay Per Crawl: should you charge AI crawlers?](https://aiscan.site/blog/cloudflare-pay-per-crawl-should-you-charge) covers the mechanics in full; the relevant point here is narrower: the same identity question this article is measuring, which IP belongs to which crawler, and how current that mapping is, is no longer only a bot-access decision. It is the input to whether a request gets charged, refused, or served for free. A stale allowlist used to mean a missed crawler. It can now mean an operator paying for traffic your published ranges say does not exist, or your own site refusing a legitimate paying crawler because the file you checked it against was nineteen months old.

[Stealth crawling vs. user-driven fetching](https://aiscan.site/blog/stealth-crawling-user-driven-fetching-debate) measured a related edge on 75 real sites: a `Disallow: /` aimed at a training crawler is backed by an actual server-side refusal only 45 to 67 percent of the time, depending on the operator. Declaring a rule and enforcing it are two different projects, and the IP files this article measures are the enforcement half for the small number of publishers who do more than write a `robots.txt` line and hope. Whether a declared block actually stops training on already-published content is a separate question again, covered in [Does blocking AI training protect your content?](https://aiscan.site/blog/does-blocking-ai-training-protect-content), but it rests on the same gap: a declaration is not a verified outcome.

## What AIScan checks, and what it still gets wrong

AIScan's [bot-access dimension](https://aiscan.site/docs/checks/bot-access) covers both checks named in this article's tags. B2 reads whether a site's `robots.txt` names any known AI agent at all, which is a one-time, static read of a file the site controls. B3 checks for a published Web Bot Auth key directory, the mechanism the previous section just described as a live payment gate. According to AIScan's own production database, queried live across 739 real sites on 26 September 2026, B2 passes on 40.1 percent and B3 on 7.7 percent, both dimensions AIScan can grade purely from what a site publishes.

What AIScan cannot see is any of the fifteen files this article fetched by hand. Verifying that a specific *inbound request* actually originated from the IP range an operator currently publishes requires access to that request's source address at the moment it arrives, which is server-log or edge-log data no external scanner is ever handed. A 100/100 AIScan score is a statement about what a site declares and how it responds to a probe, not a guarantee that its edge is checking incoming crawler traffic against a current file. Those are different questions, and conflating them is exactly the mistake this whole article is about: trusting a static signal to answer a question that only a fresh check can answer.

B3 has a bug worth naming rather than hiding, because it is one more instance of the same problem: a piece of published metadata giving no way to tell current from stale. Grouped by evidence string across the same 739-site read, B3 reports the literal string `HTTP 200` for 57 sites marked `pass` and the identical literal string `HTTP 200` for 51 sites marked `partial`. The verdict differs; the evidence given for the verdict does not. First published against a 499-site read on 9 September, still reproducing today. Anyone who reads a check's status without reading what actually backs it, ours included, is trusting a label instead of checking the file, which is precisely the habit this article is arguing against.

## How to check today, in under a minute

The fastest way to see where a specific site stands on this is to run it through AIScan: paste the URL at [aiscan.site](https://aiscan.site/) or run `npx aiscan-cli yoursite.com`, no account required, and read the B2 and B3 rows directly. That answers what the site declares and publishes.

To check an operator's IP file yourself, by hand, the command is one line per file:

```bash
curl -sL https://openai.com/gptbot.json | python3 -c "import json,sys; d=json.load(sys.stdin); print(d.get('creationTime'), len(d.get('prefixes', d.get('ips', []))))"
```

Swap in any of the fifteen URLs above. `-L` is not optional; three of them redirect. If a file's `creationTime` is months old, that is not automatically a problem, some operators genuinely do not need to change often, but it means the file is telling you what it looked like months ago, and you are the one deciding whether that is still good enough for what you are about to trust it for.

Because this article's own headline finding is that a three-week-old snapshot already misrepresented one of fifteen files, treat every number in the table above the same way: current as of 26 September 2026, not current forever. The fix is not to memorize which operators are fast and which are slow, it is to make the one-line check above part of whatever process relies on these files, on a schedule that matches how often the file in question has actually moved historically, monthly for the daily-regenerating ones is overkill, and a single annual check for a file frozen nineteen months is still a check worth doing, because nineteen months is exactly long enough for an operator to finally move it without anyone downstream noticing.

For a WordPress site making any of these declarations, [ThinkRank](https://thinkrank.ai) manages `robots.txt`, robots meta tags, schema and `llms.txt` from one plugin rather than three separate ones fighting over the same file, and it migrates existing settings from Rank Math, Yoast, All in One SEO or SEOPress, so adopting it costs nothing already configured. For a Shopify store weighing the same AI-bot-access questions, [StoreSEO](https://storeseo.com/) covers the `llms.txt` and `agents.md` half of the same problem for commerce sites specifically.

## Your next step

Run your own site through AIScan and read the B2 and B3 rows: [aiscan.site](https://aiscan.site/) or `npx aiscan-cli yoursite.com`. Then pick the one operator whose crawler matters most to your traffic and check its `creationTime` field directly, today, with the one-line command above. A file that was fresh in a previous article, including this one, is not a fact about the file, only about the day it was read. For everything else this dimension covers, the full checklist lives at [aiscan.site/guides](https://aiscan.site/guides).

