---
title: "Web Bot Auth: Cryptographic Crawler Identity, Explained"
slug: web-bot-auth-explained
published: 2026-09-26T08:21:21.717418+00:00
updated: 2026-09-26T08:21:21.717418+00:00
author: "Asif Rahman"
author_url: https://masifrahman.com
category: "AI Readiness"
tags: check:B3, web bot auth, crawler identity, http message signatures, bot verification
description: "Web Bot Auth signs HTTP requests with a private key, verified against a public JWKS directory the crawler publishes. The mechanism, and how to check yours."
url: https://aiscan.site/blog/web-bot-auth-explained
---

Across the 499 sites in our own scan corpus, exactly 19 publish a working Web Bot Auth key directory. That's 3.8%. If you run any kind of automated client that fetches other people's sites (a scraping bot, an AI assistant, a monitoring service, a shopping agent), you're one of a small minority if you can cryptographically prove who you are when you show up. Here's the mechanism, how to join that 3.8%, and how to check whether you already did it correctly.

## Quick summary

| If you want to... | Do this | What it proves | Time required |
|---|---|---|---|
| Prove your crawler's identity cryptographically | Publish a JWKS key directory at `/.well-known/http-message-signatures-directory` | The signature on a request, not just the string in its User-Agent | ~30 min with OpenSSL |
| Sign requests without hand-rolling RFC 9421 | Use an existing client (Guzzle, Linzer, Bot-Authentication, a Cloudflare Worker) | Same proof, someone else's tested signing code | ~1 afternoon of integration |
| Check whether your directory is even reachable | Scan your domain at [aiscan.site](https://aiscan.site/) for check B3 | A parseable, well-known-path JWKS response | 1 scan, free |
| Check whether your signed requests are actually well formed | `curl https://crawltest.com/cdn-cgi/web-bot-auth` with and without your signature headers | The 400 / 401 / 200 split that tells you exactly what's wrong | 1 request |

## The problem an IP list and a user-agent string can't solve

Ask any site that gets crawled by AI systems how it tells a real crawler from an impostor, and the honest answer today is: it mostly doesn't. [Reverse DNS verification only works for four operators](https://aiscan.site/blog/ai-crawler-user-agent-list-2026): Google, Apple, Bing and Common Crawl publish a hostname suffix a lookup can confirm. GPTBot, ChatGPT-User and ClaudeBot publish no such suffix, so a verifier has no method at all, and a request forged from a rented cloud IP with `ClaudeBot` in its User-Agent string passes every check a site can run against it. A user-agent string is a claim a client makes about itself. It proves nothing, and publishers already know it: of 91 publisher hosts that served a browser a real `robots.txt`, six refused a self-declared `Googlebot/2.1` outright rather than trust the header.

Web Bot Auth replaces the claim with a cryptographic proof. Instead of asking a request to identify itself with a string anyone can type, it asks the request to carry a signature over its own headers, made with a private key only the real operator holds, verifiable against a public key that operator publishes at a fixed address on its own domain. The verifier doesn't need to trust a claim. It fetches the key, checks the signature, and either the math works or it doesn't.

The specification is `draft-ietf-webbotauth-httpsig-protocol`, published 1 September 2026 by the IETF's Web Bot Auth Working Group on the standards track, authored by engineers at Cloudflare and Google. It builds on [RFC 9421, HTTP Message Signatures](https://www.rfc-editor.org/rfc/rfc9421), which already defines how to sign an HTTP request; Web Bot Auth defines what a crawler signs, where it publishes the key that verifies the signature, and how a verifier resolves one from the other.

## How the signature and the directory fit together

Three pieces do the work. First, a `Signature-Agent` header on every signed request, carrying the HTTPS URL where the operator publishes its keys. Second, an HTTP Message Signature (the `Signature` and `Signature-Input` headers from RFC 9421) computed over that URL, the request's `@authority`, and a handful of other components, using a private key. Third, a key directory at a fixed well-known path, `/.well-known/http-message-signatures-directory`, serving the matching public key as a JSON Web Key Set, the same JWKS format OAuth and OpenID Connect already use for token verification.

The directory format is deliberately plain:

```json
{
  "keys": [{
    "kty": "OKP",
    "crv": "Ed25519",
    "kid": "NFcWBst6DXG-N35nHdzMrioWntdzNZghQSkjHNMMSjw",
    "x": "JrQLj5P_89iXES9-vFgrIy29clF9CC_oPPsw3c5D0bs",
    "use": "sig",
    "nbf": 1712793600,
    "exp": 1715385600
  }]
}
```

The spec calls out Ed25519 as the algorithm both major implementers ship, though the underlying RFC also allows RSASSA-PSS. The response itself must be served over HTTPS with `Content-Type: application/http-message-signatures-directory+json`, a media type the IETF draft registers with IANA alongside the well-known path itself, so a directory served at the wrong content type fails validation even when the JSON is correct.

The identity the system produces is not the key. It's the URL. Per the spec's own framing, a verifier that resolves a `Signature-Agent` URL and confirms it publishes a key that checks out ends up with a pair: the URL, and proof that whoever controls it signed this request. That pairing is what lets an operator rotate keys without losing continuity. Publish the new key alongside the old one, then drop the old one once its window has passed, and the URL a site has been allowlisting for months never changes. It also means an operator abandoning that URL is the whole trust model's honest failure mode: nothing stops anyone from standing up a fresh domain and starting over, the same way nothing stops a spammer registering a new one today. Web Bot Auth cryptographically confirms continuity of a claimed identity. It was never designed to vouch for intent, and the spec says as much directly: a valid signature says who signed the request, not whether the request itself is authorized or benign.

Two operators have already shipped it. Cloudflare's 28 August 2026 post on [BotBase for Operators](https://blog.cloudflare.com/botbase-for-operators/) states plainly that "ChatGPT Agent is signing http requests using the newly proposed open standard Web Bot Auth," and that BotBase now validates those signatures as part of operator onboarding. Google's own crawler documentation lists a `Google-Agent` identity signing under the same protocol at `https://agent.bot.goog`. Adoption elsewhere is thin: our own count across 499 scanned sites puts the check that grades a published directory at 19 passes, or 3.8%. [We measured that, and a bug in how the check reports it, in a companion piece](https://aiscan.site/blog/stealth-crawling-user-driven-fetching-debate) rather than repeating it here.

## Who this is actually for

If your site only receives crawler traffic and never sends any, none of this applies to you yet. Web Bot Auth is something the crawler operator publishes, not something a passive website needs to add. It matters the moment your own product fetches other people's URLs: a monitoring service, a scraper, an AI browsing agent, a shopping bot doing agentic checkout, or a scanner like ours. If that's you, the sites you crawl are increasingly the ones enforcing identity at the edge, and an unsigned request is starting to look the way an unauthenticated API call looks today.

## Path 1: publish your own directory by hand

This is the free path, and it's what Cloudflare's own integration docs walk through for any operator connecting to their network.

1. **Generate an Ed25519 key pair.** `openssl genpkey -algorithm ed25519 -out private-key.pem`, then `openssl pkey -in private-key.pem -pubout -out public-key.pem`. Keep the private key off any server that isn't the one doing the signing.
2. **Convert the public key to JWK.** The [`jwker`](https://github.com/jphastings/jwker) command-line tool does this in one call: `jwker public-key.pem public-key.jwk`. Never include the `d` parameter anywhere public; that's the private half of the key.
3. **Host the directory** at `/.well-known/http-message-signatures-directory`, over HTTPS, serving `{"keys":[<your JWK>]}` with `Content-Type: application/http-message-signatures-directory+json`. The response itself should carry its own `Signature` and `Signature-Input` headers with `tag="http-message-signatures-directory"`, so nobody can mirror your directory onto another domain and claim your identity.
4. **Sign your outbound requests** with `Signature-Agent` pointing at that URL, plus `Signature-Input` naming `tag="web-bot-auth"`, a `keyid` set to the JWK thumbprint, and a short `expires` window (Cloudflare recommends about a minute), since the protocol relies on a short lifetime rather than defining nonce-based replay protection.
5. **Register the directory URL** with any network you want to be recognized by. On Cloudflare, that's Manage Account → Configurations → Bot Submission Form, with Verification Method set to Request Signature and the directory URL entered as the validation instructions.

## Path 2: sign with an existing library instead of hand-rolling it

Hand-computing an RFC 9421 signature is fiddly enough that most teams shouldn't. The protocol draft's own implementation appendix lists working clients across several languages: a Cloudflare Workers library and a Chrome extension in TypeScript, a Puppeteer script in JavaScript, a Guzzle middleware for PHP, a `Bot-Authentication` package and an HTTPie plugin in Python, and a Ruby gem called Linzer, plus integrations for the Scrapy and crawl4ai scraping frameworks. If your crawler already runs on one of those stacks, wiring in an existing signer costs an afternoon, not a spec reading.

## Verify it: AIScan first, then by hand if you'd rather

[Run your domain through AIScan](https://aiscan.site/), free, no account needed, and check the result against **check B3** under [bot access](https://aiscan.site/docs/checks/bot-access). A `pass` means the scan found a valid, parseable directory at the well-known path; `info` means nothing is published there at all, which is the default state for most sites today and not itself a failure.

If you'd rather confirm the wire format by hand, Cloudflare runs a public test endpoint for exactly this:

```bash
curl -i https://crawltest.com/cdn-cgi/web-bot-auth
```

An unsigned request returns HTTP 400 with the body `missing signature / signature-input / signature-agent headers`. We confirmed that live while researching this piece. Send a correctly formatted signature for a key nobody has registered and you get 401; a request that verifies against a key Cloudflare already knows about returns 200. That three-way split is the fastest way to tell "my headers are malformed" from "my headers are fine, this key just isn't registered anywhere yet."

## Maintenance: what changes after you publish

A directory isn't a file you write once. Rotate the signing key before the old one expires by publishing both together, then remove the old key only after its `exp` has passed. Dropping it early breaks verification for anyone whose cached copy hasn't refreshed yet. If you move your directory to a new domain, the old URL stops being your identity the moment you take it down, and every site that allowlisted it starts from zero with you again.

## What AIScan can see here, and what it can't

B3 confirms a directory exists, parses as valid JWKS, and is reachable over HTTPS. It does not confirm that the specific requests your crawler sends are actually signed, or that your `Signature-Agent` header points at the directory B3 just scanned; that only shows up in the receiving site's own logs, which no external scanner can read. Publish the directory, then check your crawler's outbound requests against `crawltest.com` directly; the two checks answer different questions, and passing one doesn't confirm the other.

## What to do next

Scan your domain at [aiscan.site](https://aiscan.site/) for B3 alongside the rest of the [bot access](https://aiscan.site/docs/checks/bot-access) checks, and if you operate a crawler of your own, work through the [guides](https://aiscan.site/guides) for publishing the other machine-readable files an agent looks for before it decides to trust your site back.

