Five identical glowing green gates in a row; the third is filled solid coral with cracks, a false pass that looks the same as the real ones from outside. AIScan.site wordmark bottom left.
Five identical glowing green gates in a row; the third is filled solid coral with cracks, a false pass that looks the same as the real ones from outside. AIScan.site wordmark bottom left.
AI Readiness

The Soft 404 Problem: Why HTTP 200 Breaks Every Agent Discovery Check

A soft 404 doesn't fail one check, it fakes every discovery file you publish. Measured on 733 sites, plus a commerce endpoint that says yes to everything.

AAsif Rahman 25 Sept 2026 18 min read
#soft 404#404 handling#agent discovery#AI readiness#RFC 9110#llms.txt

This guide covers E1 · Discoverability, C1 · Content, C2 · Content, D2 · Discoverability.

Table of contents

Every convention on the agentic web works the same way: an agent asks for a file at a well-known address, and a 200 means the file is there. /llms.txt, /robots.txt, a .md suffix on a documentation page, a UCP profile at /.well-known/ucp, an MCP server card. None of them carry a signature. None of them prove anything about the body behind the status line. The whole system runs on the assumption that a server which cannot find something says so.

Most servers keep that promise. A meaningful share do not, and the ones that break it don't fail loudly. They answer every path the same way, including the ones that were never real, and the result looks identical to genuine adoption until someone asks for a URL a second time.

Quick summary

QuestionShort answer
What is a soft 404, in one line?A response that returns 200 for a URL that doesn't exist, instead of 404
Why does it matter beyond SEO?Every agent discovery convention (llms.txt, Markdown routes, sitemaps, UCP, MCP server cards) is a probe at a well-known path, and a soft 404 makes all of them look adopted whether they are or not
How common is it?On AIScan's own scan corpus, 733 real sites, correct 404 handling (check E1) passes on 75.3%, meaning roughly one in four still fails it
Does it only break one check?No. 29.3% of the sites that pass our Markdown-negotiation check (C1) do so on a host that fails E1, so a check can pass on evidence that is itself unreliable
Can the fake side look completely convincing?Yes. Two documentation hosts return real text/markdown, not HTML, for a page that was never built, and one live commerce endpoint answers three unrelated well-known paths with the same 200
Does status alone ever mislead in the other direction?Yes. stackoverflow.com/robots.txt answers HTTP 418 while serving a completely valid, enforced file underneath
What's the fix?One extra request per file: fetch a path you know cannot exist, and compare it against the file you actually care about, on status, content type and size
Where does AIScan fit?npx aiscan-cli yoursite.com runs the control-path test as part of E1, C1, C2 and D2, and names the exact check that failed

A status code is a claim, not a fact

RFC 9110 is unusually direct about what a status code is allowed to promise. Section 12.2, fetched from the RFC Editor on 25 September 2026, describes the two status codes HTTP actually reserves for "I can't give you what you asked for": a 300 or a 406 response is supposed to carry "information about available representations so that the user or user agent can react by making a selection." That's the honest failure mode: say what you don't have, and say what exists instead.

A soft 404 skips that step entirely. Instead of a 406 or a 404, the server hands back a 200 and whatever its router had lying around, usually the same shell it hands back for every other unmatched path. Nothing in the exchange is technically false. The status line says "here is a representation," and there is one. It just isn't a representation of the thing that was asked for, and the client has no way to tell the difference from a single request.

That gap is not new. According to Search Central's guidance on HTTP errors, verified on 25 September 2026, when a page's content "suggests an error for Google Search, an empty page or an error message," Search Console records it as a soft 404 even though the wire-level status said everything was fine. What's changed since that guidance was written for search crawlers is the number of things now built on the same trust. A search index absorbing a handful of soft 404s is a ranking nuisance. An agent trying to decide whether your site publishes llms.txt, a machine-readable commerce profile, or Markdown at all is trying to answer a yes-or-no question with a signal that has no reliable "no."

Five conventions, one shared assumption

The reason this generalizes is structural rather than coincidental. Every discovery convention released since 2023 has the identical shape: put a file at a predictable address, and let a client ask for it directly rather than crawling a whole site to find it. That shape is what makes the conventions cheap to adopt. It's also exactly what a soft-404 router defeats, because the router doesn't distinguish "the requested resource" from "any resource."

ConventionAIScan checkWhat the check assumesWhat a soft 200 makes possible
Correct 404 / 410 statusE1An unmapped path returns a real error statusNothing else on this list can be trusted until this one holds
.md suffix, /index.md, or Accept: text/markdownC1A 200 with the right content type means Markdown is really being servedA catch-all shell answers 200 for /any-page.md, and a status-only checker calls it a pass
/llms.txtC2A 200 at the well-known path means the index existsA missing file and a real one are indistinguishable without reading the body
XML sitemapD2Every URL the sitemap lists resolves to real contentA sitemap is a set of claims about URLs, and nothing in the sitemap format itself re-verifies that a listed URL still resolves to anything other than the same shell

Two conventions outside this post's tagged checks show the identical failure at a larger scale. agentic-commerce-protocols-explained measured 30 live storefronts against the Universal Commerce Protocol, x402 and the MCP server card, three unrelated .well-known paths defined by three unrelated standards bodies. All three still depend on the same unglamorous prerequisite: a request for a path nobody built has to come back 404. When it doesn't, all three checks fail together, for a reason none of the three specifications anticipated.

Why even a framework's own 404 convention can leak a 200

The interesting part is that this isn't only a case of developers not bothering to wire up a 404 handler. Sometimes the framework's own documented behavior admits the leak directly. Next.js ships a dedicated not-found.js convention specifically to render a proper missing-page UI, and its own reference page, fetched with Accept: text/markdown on 25 September 2026, states plainly that "Next.js will return a 200 HTTP status code for streamed responses, and 404 for non-streamed responses." That single sentence, dated 10 July 2026 in the page's own front matter, is a framework maintainer telling you upfront that the correct status code depends on a rendering detail most teams never think about.

The reason is not carelessness. Once a server starts streaming a response, the status line has already gone out on the wire before the framework knows whether the eventual content will resolve to a real page or a not-found boundary. HTTP has no mechanism to take a status code back after the first byte. A framework that wants to stream content as fast as possible for real pages inherits a version of the same problem this whole article is about, on its own official not-found path, for a reason baked into the protocol rather than a bug in anyone's code. nextjs-react-invisible-to-ai-crawlers covers the sibling failure this same rendering pipeline produces, where a real page ships an empty body to a crawler that never runs the client-side JavaScript; this is the same mechanism, applied to the status line instead of the body.

That's also why the control-path test in this article never asks what framework a site runs. It doesn't need to know whether the 200 came from a hand-rolled catch-all route, a cache layer serving a stale page, or a streaming response that hadn't resolved yet when the headers were written. All three produce the identical externally visible signature, and all three are caught by the same two requests.

The reusable test, in two requests

The fix predates every convention it protects, because it isn't really about llms.txt or UCP. It's a property test on the router itself: fetch a path that cannot possibly exist, and compare what comes back against the file you actually care about.

# 1. A path nobody built
curl -s -o /dev/null -w '%{http_code} %{size_download} %{content_type}\n' \
  https://yoursite.com/aiscan-control-never-built-2026.md

# 2. The file you're checking
curl -s -o /dev/null -w '%{http_code} %{size_download} %{content_type}\n' \
  https://yoursite.com/llms.txt

Three outcomes, and only one of them means the file is real. If the control path returns a real 404, the site's router is honest and every other test in this article means what it says. If the control path returns 200 with a byte count close to the real file's, the router is a shell answering everything, and no check that only reads the status line can be trusted on this host. The third outcome is the one worth knowing about even though it's rare: the control path can return 200 with a real, well-formed body that isn't HTML at all, which the next section covers, because it defeats the obvious version of this same test.

aiscan-ci-gate-github-actions already built this exact test into a CI gate for llms.txt specifically, and measured it live: on a set of Replit-deployed apps, a nonexistent path and the real llms.txt path returned byte-identical shells, which meant a naive curl -f gate had been passing since the app was created. That post's version is scoped to one convention and one deployment target. What's true here is that the same two-request shape is the only thing standing between "my scanner says this passes" and "my scanner read a phantom," on any convention that lives at a well-known path, not only llms.txt.

A commerce endpoint that says yes to everything

The clearest live example of this failure isn't on a documentation site. It's a checkout flow. Verified on 25 September 2026, www.warbyparker.com was asked for three addresses defined by three different standards: an x402 payment declaration, an MPP pricing document, and an MCP server card. Every one of the three came back HTTP 200, tagged text/html, and sized within a few hundred bytes of the others, all around 184 kilobytes. None of the three protocols the paths represent is actually implemented. What's being served is one JavaScript application shell, three times, because the router doesn't know the difference between a page it has and a path it's never heard of.

A checker asserting only the status code would report this storefront as running a Universal Commerce Protocol profile, an x402 payment endpoint, and a live MCP server, three separate capabilities on three separate standards, from a single misconfigured catch-all route. AIScan's own M1 check has a version of this exact gap: it currently records the status code and stops, without asserting that the body parses as JSON or carries the field the protocol actually requires. Across the sites where the check currently applies, only 37 of 269 pass on that basis and 232 land in an informational state rather than a verified one, which is a fair reflection of how little the check can currently promise, not a comment on real-world adoption either way.

Two documentation hosts that don't fail this test the easy way

The two-request test above catches every shell-based soft 404 in one shot, because a fake file and a real one served by the same catch-all will always match on content type. documentation-sites-ai-agent-readiness found the harder case while probing 52 documentation platforms for Markdown support: two of them, Sentry's docs and n8n's docs, answer a nonexistent .md URL with a genuine text/markdown body, a few hundred bytes long, that says in plain words that the page wasn't found and links back to the site's own index.

That is good machine-readable error handling wearing the wrong status code, and it deserves to be described that way rather than lumped in with the shells. It also means the two-request test has to compare more than the content type. A checker asserting 200 plus text/markdown would count both of those hosts as serving real content at any path a client cares to invent, which is the mirror image of a soft 404: a response that tells the truth in the body and lies in the status line. Re-verifying the two large shell examples from that same sweep on 25 September 2026 shows the underlying problem hasn't moved: a nonexistent .md URL on Vercel's marketing site now returns 2,638,026 bytes of text/html at HTTP 200, up from 2,518,049 two weeks earlier, and the same request against Cursor's documentation returns 562,981 bytes, up from 155,640. Whatever generates those pages got bigger. Neither one learned to say no.

correct-404-for-ai-agents covers the matching half of this problem from the other end: once a site's status line is honest, what should the body of a real 404 actually contain for a request that isn't a browser. RFC 9457's application/problem+json shape is the answer for a client that asks nicely with an Accept header. Sentry and n8n's nonexistent-page bodies are close cousins of that idea, arrived at independently, minus the correct status code to go with them.

The mirror case: a status that lies about being fine

The soft 404 makes something missing look present. There's a symmetric failure where a real, working file gets reported through a status code that sounds like a malfunction, and a checker built to distrust anything but 200 and 404 throws the file out by mistake.

stackoverflow.com/robots.txt, verified on 25 September 2026 against the live route, answers HTTP 418, "I'm a teapot," while serving a complete, correctly formed Disallow: / policy underneath. That status was measured across a wider Stack Exchange sample earlier this month at five properties out of 75 hosts checked, all sharing the same platform, so it isn't a one-off typo on a single server. A crawler or scanner that treats any non-200, non-404 status as "couldn't read the file" discards a real policy over a joke status code from 1998's April Fools' RFC, which the site's operators have apparently kept running in production for years.

The two failures share a cause even though they point opposite directions. Neither one is really about the specific status code involved. Both are about a piece of infrastructure treating the three-digit number as if it were self-certifying, when RFC 9110 never promised that a status code and the body that follows it agree with each other by construction. The number is a claim the server is making. Whether it's telling the truth is a separate question, and answering it costs one more request than most tooling currently spends.

What our own corpus says today

Querying AIScan's production scan history directly, filtered to a single rubric version and deduplicated to one scan per host, gives the current-day version of this measurement rather than a fixed snapshot from earlier in September.

MetricResult, 733-site corpus, 25 September 2026
E1 (correct 404 handling) passes552 of 733 (75.3%)
C1 (Markdown negotiation) passes259 of 733 (35.3%)
C1 passes that sit on a host failing E176 of 259 (29.3%)
C1 passes whose own evidence string is NOT text/markdown115 of 259 (44.4%)

The last row is the one worth sitting with. On close to half of the sites our own C1 check currently marks as passing, the evidence the check recorded for its own decision doesn't actually say text/markdown, meaning the check is passing sites on a response it has already, in its own logging, identified as the wrong content type. That's not a hypothetical edge case sitting somewhere in a long tail. It's the majority failure mode of a check whose headline number looks reassuring until the underlying evidence gets read rather than trusted.

None of this is unique to AIScan's rubric. Cloudflare's own markdownNegotiation scanner probes one route (the Accept header) and has the identical structural blind spot: a status-and-content-type pass with no assertion that the body is what it claims to be. Every scanner grading a well-known-path convention inherits this exact gap unless it specifically builds the control-path test in, which is one extra request most rubrics, including AIScan's own before this finding, don't currently spend.

Fix it: run the control path on every file you publish, not once

The test is the same regardless of which convention you're checking, and it doesn't require a scanner to run.

BASE=https://yoursite.com
CTRL="$BASE/aiscan-control-$RANDOM-never-built.md"

for FILE in "/llms.txt" "/robots.txt" "/sitemap.xml" "/.well-known/mcp/server-card.json"; do
  C=$(curl -s -o /dev/null -w '%{http_code} %{size_download}' "$CTRL")
  R=$(curl -s -o /dev/null -w '%{http_code} %{size_download}' "$BASE$FILE")
  echo "control: $C  |  $FILE: $R"
done

Read the pairs, not the individual lines. A control that comes back 404 with a small, generic body and a real file that comes back 200 with a very different byte count is the honest case, and every check downstream of it means what it says. A control and a file landing on the identical status and a similar byte count means the router is answering everything the same way, and no amount of fixing the file's contents will move the score until the routing itself is fixed. That's a change to how unmatched paths are handled, not a change to what you put in llms.txt.

Where the fix actually lives depends on which of three different causes produced the soft 200, and the control-path test alone doesn't tell you which. A single-page app or a static host with a client-side router almost always needs a rewrite-rule change: the platform's catch-all fallback (a Netlify _redirects 200 rule, a Vercel rewrites entry, an Nginx try_files directive) is answering every unmatched path with the app shell by design, and the fix is scoping that fallback to actual application routes instead of every path on the domain, then letting genuinely unknown paths fall through to a real 404. A server-rendered app on a streaming framework, the Next.js case above, needs the non-streamed rendering path for its not-found boundary specifically, which for the App Router means keeping the route segment that calls notFound() outside of a loading.js boundary that would otherwise force a streamed response. A cache layer serving a stale 200, the common WordPress case, needs the cache purged or excluded for paths a 404 handler should own, not a code change at all.

On WordPress specifically, this failure mode usually isn't the router at all. It's a caching layer or a full-page cache plugin serving a stale 200 for a URL that a 404 handler would otherwise catch correctly, which the control-path test also catches, because a cached shell and a genuine miss produce the same signature. ThinkRank is worth naming here for a narrower reason than the general recommendation: because it owns robots.txt, the sitemap and llms.txt from a single settings screen rather than three separate plugins each writing their own copy, a stale-cache soft 404 shows up as one discrepancy to chase down instead of three plugins each blaming the other's file. Setup pulls the existing values straight out of whichever plugin was configured before, Rank Math, Yoast, All in One SEO or SEOPress, so nothing has to be re-typed to switch. For the daily editorial work, keyword tuning and readability scoring, Rank Math and Yoast are still the sharper tools if that's where most of the week goes; neither one ships anything like the control-path test.

Where AIScan fits, and what even this test can't see

npx aiscan-cli yoursite.com

Free, no account, and it runs a version of the control-path comparison as part of E1, then reports C1, C2 and D2 against a site that has already been confirmed to say no when it means no. Paste the URL at aiscan.site for the same report without the command line. Read E1 first: every other row on this list depends on it, and a scan that shows C1 or C2 passing while E1 fails is telling you the pass is standing on a host that can't be trusted to answer honestly.

What the scan cannot currently do, stated plainly rather than left implicit, is assert that a 200 marked as C1 actually carries text/markdown rather than a status-and-type combination that merely looks right, which the 44.4% figure above shows is a real gap rather than a theoretical one. It can't tell a Sentry-style honest error body from a real page without a human or a follow-up scan reading the content. And it can't see what a CDN or reverse proxy does to a request that never reaches the origin server at all, which is a separate class of problem from anything a router misconfiguration explains.

Make your site say no when it means no

Start with the two-request test against the file you most recently published or generated, /llms.txt if you have one, a .md route if your docs claim to serve Markdown, or the well-known path for any commerce or agent protocol your stack advertises. If the control path comes back 200, that's the finding, and it's worth fixing before anything else on this page, because every other measurement your site produces about itself, ours included, depends on the answer being 404.

Then run npx aiscan-cli yoursite.com, or scan at aiscan.site, and read E1, C1, C2 and D2 in that order. The discoverability checks are documented at /docs/checks/discoverability and the content checks at /docs/checks/content. llms-txt-validator-common-mistakes has the other eleven ways a real llms.txt file can still be wrong once the status line is honest, and nextjs-react-invisible-to-ai-crawlers covers the sibling failure where a page returns a correct status and an empty body instead of a wrong status and a full one. Every guide this programme publishes is indexed at /guides.

Frequently asked questions

AIScan reports E1 as a fail, but when I curl the exact URL myself I get a real, correctly formatted 404 page. What's wrong?

Check whether the 404 is coming from the origin or from a cache in front of it. A CDN or full-page cache can serve a stale 200 for the same URL AIScan requested moments earlier, especially right after a deploy, while your own curl a minute later hits a warm cache or the origin directly and gets the fixed response. Purge the cache for that specific path, wait for propagation, and re-scan. If the mismatch persists, run the two-request test in this article yourself against a genuinely invented path rather than the one you already fixed, since a router can serve a correct 404 for one path and a soft 200 for another depending on which rewrite rule matches first.

My scan shows C1 or C2 as passing, but an agent I tested with still can't read the file. Why would a passing check not mean the file works?

Because a pass on either check can be recorded from a status code and a content type alone, without the check confirming the body actually is what it claims to be. On AIScan's own current corpus, 44.4% of C1's passing sites carry an evidence string that isn't text/markdown in the first place, meaning the check is passing a response it has already logged as the wrong type. Fetch the URL yourself with curl -s <url> | head -c 200 and read the first two hundred bytes: if it's HTML rather than Markdown or a real llms.txt body, the check passed on the status line and nothing else.

Every one of my site's agentic-commerce well-known paths (UCP, x402, the MCP server card) started returning the same response, and I didn't change any of them. What happened?

Most likely a deploy changed your catch-all or fallback route rather than anything at those specific paths. Once a router answers every unmatched address with the same application shell, every well-known path you haven't explicitly implemented starts returning that shell too, all at once, which looks like three protocols breaking simultaneously when the actual cause is one routing change. Run the control-path test against a path you're certain was never implemented; if it now returns 200 with a body matching your other well-known paths, the router is the thing to fix, not any individual protocol.

What's the actual difference between a soft 404 and a normal 404?

A normal 404 tells a client the resource doesn't exist using the status line, which every HTTP client and every scanner checks first. A soft 404 tells the same client the resource exists, with a 200, while handing back content that has nothing to do with what was asked for, usually the same fallback page or application shell served for every other unmatched path. The practical difference is that a normal 404 is caught by anything checking status codes, while a soft 404 requires reading the body or comparing against a second, deliberately invented request to catch at all.

Does publishing an XML sitemap fix any part of the soft-404 problem?

No, and it can quietly make the problem harder to see. A sitemap is a list of URLs the site is claiming are real; nothing in the sitemap format itself re-checks that each listed URL still resolves to genuine content rather than the same catch-all shell everything else on a soft-404 site returns. AIScan's D2 check confirms the sitemap parses and returns valid entries, which is a different question from whether those entries still lead anywhere real. Fix the router's 404 behavior first; a sitemap on top of a soft-404 site is a list of URLs that all quietly resolve to the same wrong page.

How do I run the two-request test from this article without installing a scanner?

Two curl commands and nothing else. Request a path you are certain was never built on your site, and separately request the file you actually care about, both with -o /dev/null -w '%{http_code} %{size_download} %{content_type}\n' so you get the status, byte count and content type for each on one line. If the invented path returns a real 404, your router is trustworthy and the second request's result means what it says. If both come back with the same status and a similar byte count, the router is answering every path identically and the file you're checking may not exist at all.

I found a site that returns a real, readable `text/markdown` error message for a page that was never built, instead of HTML. Is that still a problem worth reporting?

It's a smaller and more forgivable problem than the usual soft 404, and it's worth being precise about which part is wrong. Sentry's and n8n's documentation sites do exactly this: a nonexistent Markdown URL returns a genuine, well-formed text/markdown body saying the page wasn't found, which is thoughtful machine-readable error handling. The only defect is the status line, which should read 404 rather than 200. A checker asserting only status and content type still gets fooled by it, so it's worth flagging, but it's a status-code bug on an otherwise honest response, not a fake file pretending to be real.

I fixed my catch-all route so unmatched paths return a real 404, but my AIScan E1 score didn't change on the next scan. What am I missing?

Check the scan's timing against your deploy and your cache, in that order. A scan that runs before the new build is live, or that hits a CDN edge still serving the previous cached response, will report the old behavior regardless of what your origin now does. Re-run the two-request test by hand against the live URL first: if curl now shows a real 404 but the AIScan report still shows a fail, force a fresh scan rather than reading a cached report, and purge any CDN cache for both the fixed path and a representative invented path before re-scanning.

Related guides