---
title: "The Soft 404 Problem: Why HTTP 200 Breaks Every Agent Discovery Check"
slug: soft-404-agent-discovery-problem
published: 2026-09-25T03:34:19.662786+00:00
updated: 2026-09-25T03:34:19.662786+00:00
author: "Asif Rahman"
author_url: https://masifrahman.com
category: "AI Readiness"
tags: check:E1, check:C1, check:C2, check:D2, soft 404, 404 handling, agent discovery, AI readiness, RFC 9110, llms.txt
description: "A soft 404 doesn't fail one check, it fakes every discovery file you publish. Measured on 733 sites, plus a commerce endpoint that says yes to everything."
url: https://aiscan.site/blog/soft-404-agent-discovery-problem
---

Every convention on the agentic web works the same way: an agent asks for a file at a well-known address, and a `200` means the file is there. `/llms.txt`, `/robots.txt`, a `.md` suffix on a documentation page, a UCP profile at `/.well-known/ucp`, an MCP server card. None of them carry a signature. None of them prove anything about the body behind the status line. The whole system runs on the assumption that a server which cannot find something says so.

Most servers keep that promise. A meaningful share do not, and the ones that break it don't fail loudly. They answer every path the same way, including the ones that were never real, and the result looks identical to genuine adoption until someone asks for a URL a second time.

## Quick summary

| Question | Short answer |
|---|---|
| What is a soft 404, in one line? | A response that returns `200` for a URL that doesn't exist, instead of `404` |
| Why does it matter beyond SEO? | Every agent discovery convention (llms.txt, Markdown routes, sitemaps, UCP, MCP server cards) is a probe at a well-known path, and a soft 404 makes all of them look adopted whether they are or not |
| How common is it? | On AIScan's own scan corpus, 733 real sites, correct 404 handling (check E1) passes on 75.3%, meaning roughly one in four still fails it |
| Does it only break one check? | No. 29.3% of the sites that pass our Markdown-negotiation check (C1) do so on a host that fails E1, so a check can pass on evidence that is itself unreliable |
| Can the fake side look completely convincing? | Yes. Two documentation hosts return real `text/markdown`, not HTML, for a page that was never built, and one live commerce endpoint answers three unrelated well-known paths with the same 200 |
| Does status alone ever mislead in the other direction? | Yes. `stackoverflow.com/robots.txt` answers `HTTP 418` while serving a completely valid, enforced file underneath |
| What's the fix? | One extra request per file: fetch a path you know cannot exist, and compare it against the file you actually care about, on status, content type and size |
| Where does AIScan fit? | `npx aiscan-cli yoursite.com` runs the control-path test as part of E1, C1, C2 and D2, and names the exact check that failed |

## A status code is a claim, not a fact

RFC 9110 is unusually direct about what a status code is allowed to promise. Section 12.2, [fetched from the RFC Editor on 25 September 2026](https://www.rfc-editor.org/rfc/rfc9110.html#name-reactive-negotiation), describes the two status codes HTTP actually reserves for "I can't give you what you asked for": a `300` or a `406` response is supposed to carry "information about available representations so that the user or user agent can react by making a selection." That's the honest failure mode: say what you don't have, and say what exists instead.

A soft 404 skips that step entirely. Instead of a `406` or a `404`, the server hands back a `200` and whatever its router had lying around, usually the same shell it hands back for every other unmatched path. Nothing in the exchange is technically false. The status line says "here is a representation," and there is one. It just isn't a representation of the thing that was asked for, and the client has no way to tell the difference from a single request.

That gap is not new. According to [Search Central's guidance on HTTP errors](https://developers.google.com/search/docs/crawling-indexing/http-network-errors), verified on 25 September 2026, when a page's content "suggests an error for Google Search, an empty page or an error message," Search Console records it as a soft 404 even though the wire-level status said everything was fine. What's changed since that guidance was written for search crawlers is the number of things now built on the same trust. A search index absorbing a handful of soft 404s is a ranking nuisance. An agent trying to decide whether your site publishes llms.txt, a machine-readable commerce profile, or Markdown at all is trying to answer a yes-or-no question with a signal that has no reliable "no."

## Five conventions, one shared assumption

The reason this generalizes is structural rather than coincidental. Every discovery convention released since 2023 has the identical shape: put a file at a predictable address, and let a client ask for it directly rather than crawling a whole site to find it. That shape is what makes the conventions cheap to adopt. It's also exactly what a soft-404 router defeats, because the router doesn't distinguish "the requested resource" from "any resource."

| Convention | AIScan check | What the check assumes | What a soft 200 makes possible |
|---|---|---|---|
| Correct 404 / 410 status | E1 | An unmapped path returns a real error status | Nothing else on this list can be trusted until this one holds |
| `.md` suffix, `/index.md`, or `Accept: text/markdown` | C1 | A 200 with the right content type means Markdown is really being served | A catch-all shell answers `200` for `/any-page.md`, and a status-only checker calls it a pass |
| `/llms.txt` | C2 | A 200 at the well-known path means the index exists | A missing file and a real one are indistinguishable without reading the body |
| XML sitemap | D2 | Every URL the sitemap lists resolves to real content | A sitemap is a set of claims about URLs, and nothing in the sitemap format itself re-verifies that a listed URL still resolves to anything other than the same shell |

Two conventions outside this post's tagged checks show the identical failure at a larger scale. [`agentic-commerce-protocols-explained`](https://aiscan.site/blog/agentic-commerce-protocols-explained) measured 30 live storefronts against the Universal Commerce Protocol, x402 and the MCP server card, three unrelated `.well-known` paths defined by three unrelated standards bodies. All three still depend on the same unglamorous prerequisite: a request for a path nobody built has to come back `404`. When it doesn't, all three checks fail together, for a reason none of the three specifications anticipated.

## Why even a framework's own 404 convention can leak a 200

The interesting part is that this isn't only a case of developers not bothering to wire up a 404 handler. Sometimes the framework's own documented behavior admits the leak directly. Next.js ships a dedicated `not-found.js` convention specifically to render a proper missing-page UI, and its own reference page, [fetched with `Accept: text/markdown` on 25 September 2026](https://nextjs.org/docs/app/api-reference/file-conventions/not-found), states plainly that "Next.js will return a `200` HTTP status code for streamed responses, and `404` for non-streamed responses." That single sentence, dated 10 July 2026 in the page's own front matter, is a framework maintainer telling you upfront that the correct status code depends on a rendering detail most teams never think about.

The reason is not carelessness. Once a server starts streaming a response, the status line has already gone out on the wire before the framework knows whether the eventual content will resolve to a real page or a not-found boundary. HTTP has no mechanism to take a status code back after the first byte. A framework that wants to stream content as fast as possible for real pages inherits a version of the same problem this whole article is about, on its own official not-found path, for a reason baked into the protocol rather than a bug in anyone's code. [`nextjs-react-invisible-to-ai-crawlers`](https://aiscan.site/blog/nextjs-react-invisible-to-ai-crawlers) covers the sibling failure this same rendering pipeline produces, where a real page ships an empty body to a crawler that never runs the client-side JavaScript; this is the same mechanism, applied to the status line instead of the body.

That's also why the control-path test in this article never asks what framework a site runs. It doesn't need to know whether the 200 came from a hand-rolled catch-all route, a cache layer serving a stale page, or a streaming response that hadn't resolved yet when the headers were written. All three produce the identical externally visible signature, and all three are caught by the same two requests.

## The reusable test, in two requests

The fix predates every convention it protects, because it isn't really about llms.txt or UCP. It's a property test on the router itself: fetch a path that cannot possibly exist, and compare what comes back against the file you actually care about.

```bash
# 1. A path nobody built
curl -s -o /dev/null -w '%{http_code} %{size_download} %{content_type}\n' \
  https://yoursite.com/aiscan-control-never-built-2026.md

# 2. The file you're checking
curl -s -o /dev/null -w '%{http_code} %{size_download} %{content_type}\n' \
  https://yoursite.com/llms.txt
```

Three outcomes, and only one of them means the file is real. If the control path returns a real `404`, the site's router is honest and every other test in this article means what it says. If the control path returns `200` with a byte count close to the real file's, the router is a shell answering everything, and no check that only reads the status line can be trusted on this host. The third outcome is the one worth knowing about even though it's rare: the control path can return `200` with a real, well-formed body that isn't HTML at all, which the next section covers, because it defeats the obvious version of this same test.

[`aiscan-ci-gate-github-actions`](https://aiscan.site/blog/aiscan-ci-gate-github-actions) already built this exact test into a CI gate for llms.txt specifically, and measured it live: on a set of Replit-deployed apps, a nonexistent path and the real llms.txt path returned byte-identical shells, which meant a naive `curl -f` gate had been passing since the app was created. That post's version is scoped to one convention and one deployment target. What's true here is that the same two-request shape is the only thing standing between "my scanner says this passes" and "my scanner read a phantom," on any convention that lives at a well-known path, not only llms.txt.

## A commerce endpoint that says yes to everything

The clearest live example of this failure isn't on a documentation site. It's a checkout flow. Verified on 25 September 2026, `www.warbyparker.com` was asked for three addresses defined by three different standards: an x402 payment declaration, an MPP pricing document, and an MCP server card. Every one of the three came back `HTTP 200`, tagged `text/html`, and sized within a few hundred bytes of the others, all around 184 kilobytes. None of the three protocols the paths represent is actually implemented. What's being served is one JavaScript application shell, three times, because the router doesn't know the difference between a page it has and a path it's never heard of.

A checker asserting only the status code would report this storefront as running a Universal Commerce Protocol profile, an x402 payment endpoint, and a live MCP server, three separate capabilities on three separate standards, from a single misconfigured catch-all route. AIScan's own M1 check has a version of this exact gap: it currently records the status code and stops, without asserting that the body parses as JSON or carries the field the protocol actually requires. Across the sites where the check currently applies, only 37 of 269 pass on that basis and 232 land in an informational state rather than a verified one, which is a fair reflection of how little the check can currently promise, not a comment on real-world adoption either way.

## Two documentation hosts that don't fail this test the easy way

The two-request test above catches every shell-based soft 404 in one shot, because a fake file and a real one served by the same catch-all will always match on content type. [`documentation-sites-ai-agent-readiness`](https://aiscan.site/blog/documentation-sites-ai-agent-readiness) found the harder case while probing 52 documentation platforms for Markdown support: two of them, Sentry's docs and n8n's docs, answer a nonexistent `.md` URL with a genuine `text/markdown` body, a few hundred bytes long, that says in plain words that the page wasn't found and links back to the site's own index.

That is good machine-readable error handling wearing the wrong status code, and it deserves to be described that way rather than lumped in with the shells. It also means the two-request test has to compare more than the content type. A checker asserting `200` plus `text/markdown` would count both of those hosts as serving real content at any path a client cares to invent, which is the mirror image of a soft 404: a response that tells the truth in the body and lies in the status line. Re-verifying the two large shell examples from that same sweep on 25 September 2026 shows the underlying problem hasn't moved: a nonexistent `.md` URL on Vercel's marketing site now returns 2,638,026 bytes of `text/html` at `HTTP 200`, up from 2,518,049 two weeks earlier, and the same request against Cursor's documentation returns 562,981 bytes, up from 155,640. Whatever generates those pages got bigger. Neither one learned to say no.

[`correct-404-for-ai-agents`](https://aiscan.site/blog/correct-404-for-ai-agents) covers the matching half of this problem from the other end: once a site's status line is honest, what should the body of a real 404 actually contain for a request that isn't a browser. RFC 9457's `application/problem+json` shape is the answer for a client that asks nicely with an `Accept` header. Sentry and n8n's nonexistent-page bodies are close cousins of that idea, arrived at independently, minus the correct status code to go with them.

## The mirror case: a status that lies about being fine

The soft 404 makes something missing look present. There's a symmetric failure where a real, working file gets reported through a status code that sounds like a malfunction, and a checker built to distrust anything but `200` and `404` throws the file out by mistake.

`stackoverflow.com/robots.txt`, verified on 25 September 2026 against the live route, answers `HTTP 418`, "I'm a teapot," while serving a complete, correctly formed `Disallow: /` policy underneath. That status was measured across a wider Stack Exchange sample earlier this month at five properties out of 75 hosts checked, all sharing the same platform, so it isn't a one-off typo on a single server. A crawler or scanner that treats any non-200, non-404 status as "couldn't read the file" discards a real policy over a joke status code from 1998's April Fools' RFC, which the site's operators have apparently kept running in production for years.

The two failures share a cause even though they point opposite directions. Neither one is really about the specific status code involved. Both are about a piece of infrastructure treating the three-digit number as if it were self-certifying, when RFC 9110 never promised that a status code and the body that follows it agree with each other by construction. The number is a claim the server is making. Whether it's telling the truth is a separate question, and answering it costs one more request than most tooling currently spends.

## What our own corpus says today

Querying AIScan's production scan history directly, filtered to a single rubric version and deduplicated to one scan per host, gives the current-day version of this measurement rather than a fixed snapshot from earlier in September.

| Metric | Result, 733-site corpus, 25 September 2026 |
|---|---|
| E1 (correct 404 handling) passes | 552 of 733 (75.3%) |
| C1 (Markdown negotiation) passes | 259 of 733 (35.3%) |
| C1 passes that sit on a host failing E1 | 76 of 259 (29.3%) |
| C1 passes whose own evidence string is NOT `text/markdown` | 115 of 259 (44.4%) |

The last row is the one worth sitting with. On close to half of the sites our own C1 check currently marks as passing, the evidence the check recorded for its own decision doesn't actually say `text/markdown`, meaning the check is passing sites on a response it has already, in its own logging, identified as the wrong content type. That's not a hypothetical edge case sitting somewhere in a long tail. It's the majority failure mode of a check whose headline number looks reassuring until the underlying evidence gets read rather than trusted.

None of this is unique to AIScan's rubric. Cloudflare's own `markdownNegotiation` scanner probes one route (the `Accept` header) and has the identical structural blind spot: a status-and-content-type pass with no assertion that the body is what it claims to be. Every scanner grading a well-known-path convention inherits this exact gap unless it specifically builds the control-path test in, which is one extra request most rubrics, including AIScan's own before this finding, don't currently spend.

## Fix it: run the control path on every file you publish, not once

The test is the same regardless of which convention you're checking, and it doesn't require a scanner to run.

```bash
BASE=https://yoursite.com
CTRL="$BASE/aiscan-control-$RANDOM-never-built.md"

for FILE in "/llms.txt" "/robots.txt" "/sitemap.xml" "/.well-known/mcp/server-card.json"; do
  C=$(curl -s -o /dev/null -w '%{http_code} %{size_download}' "$CTRL")
  R=$(curl -s -o /dev/null -w '%{http_code} %{size_download}' "$BASE$FILE")
  echo "control: $C  |  $FILE: $R"
done
```

Read the pairs, not the individual lines. A control that comes back `404` with a small, generic body and a real file that comes back `200` with a very different byte count is the honest case, and every check downstream of it means what it says. A control and a file landing on the identical status and a similar byte count means the router is answering everything the same way, and no amount of fixing the file's contents will move the score until the routing itself is fixed. That's a change to how unmatched paths are handled, not a change to what you put in llms.txt.

Where the fix actually lives depends on which of three different causes produced the soft 200, and the control-path test alone doesn't tell you which. A single-page app or a static host with a client-side router almost always needs a rewrite-rule change: the platform's catch-all fallback (a Netlify `_redirects` `200` rule, a Vercel `rewrites` entry, an Nginx `try_files` directive) is answering every unmatched path with the app shell by design, and the fix is scoping that fallback to actual application routes instead of every path on the domain, then letting genuinely unknown paths fall through to a real 404. A server-rendered app on a streaming framework, the Next.js case above, needs the non-streamed rendering path for its not-found boundary specifically, which for the App Router means keeping the route segment that calls `notFound()` outside of a `loading.js` boundary that would otherwise force a streamed response. A cache layer serving a stale 200, the common WordPress case, needs the cache purged or excluded for paths a 404 handler should own, not a code change at all.

On WordPress specifically, this failure mode usually isn't the router at all. It's a caching layer or a full-page cache plugin serving a stale `200` for a URL that a 404 handler would otherwise catch correctly, which the control-path test also catches, because a cached shell and a genuine miss produce the same signature. [ThinkRank](https://thinkrank.ai) is worth naming here for a narrower reason than the general recommendation: because it owns robots.txt, the sitemap and llms.txt from a single settings screen rather than three separate plugins each writing their own copy, a stale-cache soft 404 shows up as one discrepancy to chase down instead of three plugins each blaming the other's file. Setup pulls the existing values straight out of whichever plugin was configured before, Rank Math, Yoast, All in One SEO or SEOPress, so nothing has to be re-typed to switch. For the daily editorial work, keyword tuning and readability scoring, Rank Math and Yoast are still the sharper tools if that's where most of the week goes; neither one ships anything like the control-path test.

## Where AIScan fits, and what even this test can't see

```bash
npx aiscan-cli yoursite.com
```

Free, no account, and it runs a version of the control-path comparison as part of E1, then reports C1, C2 and D2 against a site that has already been confirmed to say no when it means no. Paste the URL at [aiscan.site](https://aiscan.site/) for the same report without the command line. Read E1 first: every other row on this list depends on it, and a scan that shows C1 or C2 passing while E1 fails is telling you the pass is standing on a host that can't be trusted to answer honestly.

What the scan cannot currently do, stated plainly rather than left implicit, is assert that a `200` marked as C1 actually carries `text/markdown` rather than a status-and-type combination that merely looks right, which the 44.4% figure above shows is a real gap rather than a theoretical one. It can't tell a Sentry-style honest error body from a real page without a human or a follow-up scan reading the content. And it can't see what a CDN or reverse proxy does to a request that never reaches the origin server at all, which is a separate class of problem from anything a router misconfiguration explains.

## Make your site say no when it means no

Start with the two-request test against the file you most recently published or generated, `/llms.txt` if you have one, a `.md` route if your docs claim to serve Markdown, or the well-known path for any commerce or agent protocol your stack advertises. If the control path comes back `200`, that's the finding, and it's worth fixing before anything else on this page, because every other measurement your site produces about itself, ours included, depends on the answer being `404`.

Then run `npx aiscan-cli yoursite.com`, or scan at [aiscan.site](https://aiscan.site/), and read E1, C1, C2 and D2 in that order. The discoverability checks are documented at [`/docs/checks/discoverability`](https://aiscan.site/docs/checks/discoverability) and the content checks at [`/docs/checks/content`](https://aiscan.site/docs/checks/content). [`llms-txt-validator-common-mistakes`](https://aiscan.site/blog/llms-txt-validator-common-mistakes) has the other eleven ways a real llms.txt file can still be wrong once the status line is honest, and [`nextjs-react-invisible-to-ai-crawlers`](https://aiscan.site/blog/nextjs-react-invisible-to-ai-crawlers) covers the sibling failure where a page returns a correct status and an empty body instead of a wrong status and a full one. Every guide this programme publishes is indexed at [`/guides`](https://aiscan.site/guides).

