Textless dark green cover illustration for an AI Scan blog post: three linked cryptographic key icons forming a triangle, one marked with a check, on a deep emerald gradient background.
Textless dark green cover illustration for an AI Scan blog post: three linked cryptographic key icons forming a triangle, one marked with a check, on a deep emerald gradient background.
AI Readiness

Web Bot Auth: Cryptographic Crawler Identity, Explained

Web Bot Auth signs HTTP requests with a private key, verified against a public JWKS directory the crawler publishes. The mechanism, and how to check yours.

AAsif Rahman 26 Sept 2026 9 min read
#web bot auth#crawler identity#http message signatures#bot verification

This guide covers B3 · Bot Access.

Table of contents

Across the 499 sites in our own scan corpus, exactly 19 publish a working Web Bot Auth key directory. That's 3.8%. If you run any kind of automated client that fetches other people's sites (a scraping bot, an AI assistant, a monitoring service, a shopping agent), you're one of a small minority if you can cryptographically prove who you are when you show up. Here's the mechanism, how to join that 3.8%, and how to check whether you already did it correctly.

Quick summary

If you want to...Do thisWhat it provesTime required
Prove your crawler's identity cryptographicallyPublish a JWKS key directory at /.well-known/http-message-signatures-directoryThe signature on a request, not just the string in its User-Agent~30 min with OpenSSL
Sign requests without hand-rolling RFC 9421Use an existing client (Guzzle, Linzer, Bot-Authentication, a Cloudflare Worker)Same proof, someone else's tested signing code~1 afternoon of integration
Check whether your directory is even reachableScan your domain at aiscan.site for check B3A parseable, well-known-path JWKS response1 scan, free
Check whether your signed requests are actually well formedcurl https://crawltest.com/cdn-cgi/web-bot-auth with and without your signature headersThe 400 / 401 / 200 split that tells you exactly what's wrong1 request

The problem an IP list and a user-agent string can't solve

Ask any site that gets crawled by AI systems how it tells a real crawler from an impostor, and the honest answer today is: it mostly doesn't. Reverse DNS verification only works for four operators: Google, Apple, Bing and Common Crawl publish a hostname suffix a lookup can confirm. GPTBot, ChatGPT-User and ClaudeBot publish no such suffix, so a verifier has no method at all, and a request forged from a rented cloud IP with ClaudeBot in its User-Agent string passes every check a site can run against it. A user-agent string is a claim a client makes about itself. It proves nothing, and publishers already know it: of 91 publisher hosts that served a browser a real robots.txt, six refused a self-declared Googlebot/2.1 outright rather than trust the header.

Web Bot Auth replaces the claim with a cryptographic proof. Instead of asking a request to identify itself with a string anyone can type, it asks the request to carry a signature over its own headers, made with a private key only the real operator holds, verifiable against a public key that operator publishes at a fixed address on its own domain. The verifier doesn't need to trust a claim. It fetches the key, checks the signature, and either the math works or it doesn't.

The specification is draft-ietf-webbotauth-httpsig-protocol, published 1 September 2026 by the IETF's Web Bot Auth Working Group on the standards track, authored by engineers at Cloudflare and Google. It builds on RFC 9421, HTTP Message Signatures, which already defines how to sign an HTTP request; Web Bot Auth defines what a crawler signs, where it publishes the key that verifies the signature, and how a verifier resolves one from the other.

How the signature and the directory fit together

Three pieces do the work. First, a Signature-Agent header on every signed request, carrying the HTTPS URL where the operator publishes its keys. Second, an HTTP Message Signature (the Signature and Signature-Input headers from RFC 9421) computed over that URL, the request's @authority, and a handful of other components, using a private key. Third, a key directory at a fixed well-known path, /.well-known/http-message-signatures-directory, serving the matching public key as a JSON Web Key Set, the same JWKS format OAuth and OpenID Connect already use for token verification.

The directory format is deliberately plain:

{
  "keys": [{
    "kty": "OKP",
    "crv": "Ed25519",
    "kid": "NFcWBst6DXG-N35nHdzMrioWntdzNZghQSkjHNMMSjw",
    "x": "JrQLj5P_89iXES9-vFgrIy29clF9CC_oPPsw3c5D0bs",
    "use": "sig",
    "nbf": 1712793600,
    "exp": 1715385600
  }]
}

The spec calls out Ed25519 as the algorithm both major implementers ship, though the underlying RFC also allows RSASSA-PSS. The response itself must be served over HTTPS with Content-Type: application/http-message-signatures-directory+json, a media type the IETF draft registers with IANA alongside the well-known path itself, so a directory served at the wrong content type fails validation even when the JSON is correct.

The identity the system produces is not the key. It's the URL. Per the spec's own framing, a verifier that resolves a Signature-Agent URL and confirms it publishes a key that checks out ends up with a pair: the URL, and proof that whoever controls it signed this request. That pairing is what lets an operator rotate keys without losing continuity. Publish the new key alongside the old one, then drop the old one once its window has passed, and the URL a site has been allowlisting for months never changes. It also means an operator abandoning that URL is the whole trust model's honest failure mode: nothing stops anyone from standing up a fresh domain and starting over, the same way nothing stops a spammer registering a new one today. Web Bot Auth cryptographically confirms continuity of a claimed identity. It was never designed to vouch for intent, and the spec says as much directly: a valid signature says who signed the request, not whether the request itself is authorized or benign.

Two operators have already shipped it. Cloudflare's 28 August 2026 post on BotBase for Operators states plainly that "ChatGPT Agent is signing http requests using the newly proposed open standard Web Bot Auth," and that BotBase now validates those signatures as part of operator onboarding. Google's own crawler documentation lists a Google-Agent identity signing under the same protocol at https://agent.bot.goog. Adoption elsewhere is thin: our own count across 499 scanned sites puts the check that grades a published directory at 19 passes, or 3.8%. We measured that, and a bug in how the check reports it, in a companion piece rather than repeating it here.

Who this is actually for

If your site only receives crawler traffic and never sends any, none of this applies to you yet. Web Bot Auth is something the crawler operator publishes, not something a passive website needs to add. It matters the moment your own product fetches other people's URLs: a monitoring service, a scraper, an AI browsing agent, a shopping bot doing agentic checkout, or a scanner like ours. If that's you, the sites you crawl are increasingly the ones enforcing identity at the edge, and an unsigned request is starting to look the way an unauthenticated API call looks today.

Path 1: publish your own directory by hand

This is the free path, and it's what Cloudflare's own integration docs walk through for any operator connecting to their network.

  1. Generate an Ed25519 key pair. openssl genpkey -algorithm ed25519 -out private-key.pem, then openssl pkey -in private-key.pem -pubout -out public-key.pem. Keep the private key off any server that isn't the one doing the signing.
  2. Convert the public key to JWK. The jwker command-line tool does this in one call: jwker public-key.pem public-key.jwk. Never include the d parameter anywhere public; that's the private half of the key.
  3. Host the directory at /.well-known/http-message-signatures-directory, over HTTPS, serving {"keys":[<your JWK>]} with Content-Type: application/http-message-signatures-directory+json. The response itself should carry its own Signature and Signature-Input headers with tag="http-message-signatures-directory", so nobody can mirror your directory onto another domain and claim your identity.
  4. Sign your outbound requests with Signature-Agent pointing at that URL, plus Signature-Input naming tag="web-bot-auth", a keyid set to the JWK thumbprint, and a short expires window (Cloudflare recommends about a minute), since the protocol relies on a short lifetime rather than defining nonce-based replay protection.
  5. Register the directory URL with any network you want to be recognized by. On Cloudflare, that's Manage Account → Configurations → Bot Submission Form, with Verification Method set to Request Signature and the directory URL entered as the validation instructions.

Path 2: sign with an existing library instead of hand-rolling it

Hand-computing an RFC 9421 signature is fiddly enough that most teams shouldn't. The protocol draft's own implementation appendix lists working clients across several languages: a Cloudflare Workers library and a Chrome extension in TypeScript, a Puppeteer script in JavaScript, a Guzzle middleware for PHP, a Bot-Authentication package and an HTTPie plugin in Python, and a Ruby gem called Linzer, plus integrations for the Scrapy and crawl4ai scraping frameworks. If your crawler already runs on one of those stacks, wiring in an existing signer costs an afternoon, not a spec reading.

Verify it: AIScan first, then by hand if you'd rather

Run your domain through AIScan, free, no account needed, and check the result against check B3 under bot access. A pass means the scan found a valid, parseable directory at the well-known path; info means nothing is published there at all, which is the default state for most sites today and not itself a failure.

If you'd rather confirm the wire format by hand, Cloudflare runs a public test endpoint for exactly this:

curl -i https://crawltest.com/cdn-cgi/web-bot-auth

An unsigned request returns HTTP 400 with the body missing signature / signature-input / signature-agent headers. We confirmed that live while researching this piece. Send a correctly formatted signature for a key nobody has registered and you get 401; a request that verifies against a key Cloudflare already knows about returns 200. That three-way split is the fastest way to tell "my headers are malformed" from "my headers are fine, this key just isn't registered anywhere yet."

Maintenance: what changes after you publish

A directory isn't a file you write once. Rotate the signing key before the old one expires by publishing both together, then remove the old key only after its exp has passed. Dropping it early breaks verification for anyone whose cached copy hasn't refreshed yet. If you move your directory to a new domain, the old URL stops being your identity the moment you take it down, and every site that allowlisted it starts from zero with you again.

What AIScan can see here, and what it can't

B3 confirms a directory exists, parses as valid JWKS, and is reachable over HTTPS. It does not confirm that the specific requests your crawler sends are actually signed, or that your Signature-Agent header points at the directory B3 just scanned; that only shows up in the receiving site's own logs, which no external scanner can read. Publish the directory, then check your crawler's outbound requests against crawltest.com directly; the two checks answer different questions, and passing one doesn't confirm the other.

What to do next

Scan your domain at aiscan.site for B3 alongside the rest of the bot access checks, and if you operate a crawler of your own, work through the guides for publishing the other machine-readable files an agent looks for before it decides to trust your site back.

Frequently asked questions

Does Web Bot Auth replace robots.txt?

No. robots.txt states what a crawler is allowed to fetch; Web Bot Auth proves which crawler is actually making the request. A site can use either alone or both together, and most of the enforcement value comes from combining a stated policy with a verifiable identity.

My key directory returns a 404 at the well-known path. What's wrong?

The path must be exactly /.well-known/http-message-signatures-directory, served over HTTPS, with no redirect. A 404 usually means the route was never registered on the server, or it's serving from the wrong subdomain than the one your Signature-Agent header points to.

I sent a signed request to crawltest.com and got HTTP 400. What does that mean?

A 400 from https://crawltest.com/cdn-cgi/web-bot-auth means the request is missing one of the three required headers: Signature, Signature-Input, or Signature-Agent. Check that all three are present and that Signature-Agent is a quoted HTTPS URL, not a bare string.

My signed request returns HTTP 401 from the test endpoint. Is my signature wrong?

Not necessarily. A 401 means the message was formatted correctly but the key it references isn't registered with that verifier yet, or the signature failed to validate against a key that is registered. Confirm the key in your directory matches the one you signed with, then register the directory URL with the network before retesting.

Which signature algorithm should I use, Ed25519 or RSASSA-PSS?

Use Ed25519 unless you have a specific reason not to. It's the algorithm both Cloudflare and Google document and ship support for, and it produces a much shorter key and signature than RSASSA-PSS, which the underlying HTTP Message Signatures RFC also permits but neither major verifier prioritizes.

Do I need this if I only run a normal website with no crawler of my own?

No. Web Bot Auth is published by the operator of an automated client, not by a passive website. If nothing you run fetches other people's URLs, there's nothing for you to publish; check bot-access checks B1 and B2 instead, which grade how your own robots.txt addresses other crawlers.

How often do I need to rotate my signing key?

The spec doesn't set a fixed interval. Rotate whenever your own security policy requires it, or immediately if a key is suspected compromised, by publishing the new key alongside the old one and removing the old one only after its exp timestamp passes.

AIScan scanned my domain and check B3 came back info instead of pass. Did I do something wrong?

Not necessarily. info means no key directory was found at the well-known path, which is the default and expected result for any site that doesn't operate its own crawler. It only becomes worth fixing if you do run an automated client and intended to publish one.

Related guides