Table of contents
An agent that lands on your homepage has already paid for the request. If that response can also say where your llms.txt, sitemap and API catalog live, the agent skips the guessing entirely. The Link: HTTP header does that job, it costs one line of server config, and most sites still leave it empty. On 29 September 2026 we fetched the homepage headers of 15 domains and only 4 sent an api-catalog relation, ours included.
Quick summary
| If you want to… | Do this | Time | What it changes |
|---|---|---|---|
| Pass check D3 | Send Link: </.well-known/api-catalog>; rel="api-catalog" on the homepage response | 5 min | The scan can find a machine-readable description without guessing |
| Point agents at your llms.txt | Add </llms.txt>; rel="describedby"; type="text/markdown" | 5 min | The header names the file, no path convention needed |
| Advertise your sitemap | Add </sitemap.xml>; rel="sitemap"; type="application/xml" | 5 min | Same idea for crawlers that read headers before bodies |
| Check what you send today | Run npx aiscan-cli yoursite.com and read D3 | 1 min | Shows the header your CDN actually delivers |
Short answer. Add one Link: header to your homepage response, listing the relations you really have. Set it at the server or CDN, then confirm with a scan. If you have no API and no llms.txt, D3 is informational and you can skip it.
Why a header, when you already have files
Every discovery file has the same weakness: the agent has to know the path first. /llms.txt is a convention, /.well-known/api-catalog is a registered path, and a sitemap can live anywhere unless robots.txt names it. The agent tries a list of guesses and gives up on the misses.
Think of the files as rooms in a building and the Link: header as the sign in the lobby. The rooms exist either way. The sign saves the visitor from trying every door.
The format comes from RFC 8288, Web Linking. Its ABNF is Link = #link-value, where each link-value is a URI reference in angle brackets followed by semicolon-separated parameters. According to RFC 8288, its own examples show two equivalent ways to send several links: one header with comma-separated values, or several Link: header lines. A recipient has to treat both the same. The rel parameter names the relationship, and per the RFC a single rel may carry more than one relation type, separated by spaces.
Which relations to send
Three relations cover what most sites have:
| Relation | Points at | Registered by | AIScan check |
|---|---|---|---|
api-catalog | /.well-known/api-catalog | RFC 9727 | D3, and P1 for the file itself |
describedby | your llms.txt | used by llms.txt v2 | D3 |
sitemap | your sitemap | de facto | D3 |
According to its own docs, AIScan grades D3 as inspecting "the HTTP Link response header on your homepage," where "a rel=api-catalog or rel=describedby entry points agents at a machine-readable description of what your site exposes." Its suggested fix is a single line: Link: </.well-known/api-catalog>; rel="api-catalog". The same page says D3 is informational on content sites without an API catalog. Publish the catalog first if you have an API, using our guide to publishing an API catalog, because a header that points at a 404 is worse than no header.
Do not add a relation for a file you do not serve. Each target in the header must return 200 to a plain GET.
Path one: nginx
- Open the server block for your site, usually a file under
/etc/nginx/sites-enabled/. - Inside
location = / { ... }add oneadd_headerline per relation, withalwaysso the header also goes out on redirects and error codes:
add_header Link '</.well-known/api-catalog>; rel="api-catalog"; type="application/linkset+json"' always;
add_header Link '</llms.txt>; rel="describedby"; type="text/markdown"' always;
add_header Link '</sitemap.xml>; rel="sitemap"; type="application/xml"' always;
- Run
nginx -t, thennginx -s reload. - Check that the header comes back with
curl -sI https://yoursite.com/ | grep -i '^link:'.
According to the nginx documentation, one trap catches people. add_header directives are inherited from the previous configuration level only if there are none at the current level. Put a single add_header X-Frame-Options ... in your location block and the server-level Link lines vanish. Repeat the whole set at the level that serves the homepage, or use add_header_inherit merge;, which the docs list from nginx 1.29.3.
Path two: Apache
- Confirm
mod_headersis enabled (apachectl -M | grep headers). - In the virtual host, or in
.htaccessif you cannot edit the host, add:
Header always set Link "</.well-known/api-catalog>; rel=\"api-catalog\"; type=\"application/linkset+json\", </llms.txt>; rel=\"describedby\"; type=\"text/markdown\""
- Reload with
apachectl graceful. - Run the same
curl -sIcheck.
set replaces any existing Link header, so if WordPress or another layer already sends one, use append for your entries instead. According to the Apache mod_headers documentation, append adds to an existing header with a comma separator, which matches the RFC 8288 list form.
Path three: a CDN rule
If you would rather not touch the origin, Cloudflare's Transform Rules have a response-header rule that adds a static header to matching responses. Use the add operation rather than set when the origin may already send a Link: header, so both survive. Match the rule to the hostname and the path /, then apply the same three values. The Cloudflare docs on modifying response headers cover the dashboard steps for your plan. Netlify and Vercel read headers from a config file in the project instead, documented at Netlify's headers page and Vercel's project configuration.
What the header should not become
A Link: header is not a place to dump every hint. Browsers use the same header for rel=preload and rel=preconnect, and it is easy to end up with a header that is all performance hints. We saw exactly that:
| Site (29 Sep 2026) | Discovery relations sent | Other relations |
|---|---|---|
| aiscan.site | api-catalog, describedby, sitemap, service-desc | none |
| cloudflare.com | api-catalog, sitemap, service-desc (3), service-doc | preload (6), preconnect (4) |
| vercel.com | api-catalog, ai-catalog, agent-skills | none |
| netlify.com | api-catalog | none |
| wordpress.org | none | REST API, alternate, shortlink |
| nextjs.org | none | preload (8) |
| wix.com | none | preconnect only |
| github.com, stripe.com, shopify.com, developer.mozilla.org, ghost.org, wikipedia.org, storeseo.com, thinkrank.ai | none | no Link header on the homepage response we fetched |
We fetched these headers from each homepage on 29 Sep 2026 with a browser user agent, one request per site, so treat them as a snapshot and not a survey. The pattern still holds: sites that send a Link: header often spend it on preloads, and the ones that declare discovery relations are rare. Wix, which sends the header on the page we fetched, uses all of it on preconnect and declares no discovery relation at all.
Verify it
Start with the scan. npx aiscan-cli yoursite.com or paste the URL at aiscan.site is free and needs no account. Read three rows: D3 (Link header for discovery), P1 (does the api-catalog target resolve) and C2 (does the llms.txt it names exist). D3 passes on the header alone. The other two grade the files behind it, so a header can pass D3 while pointing at something broken.
If you would rather check by hand:
curl -sI https://yoursite.com/ | grep -i '^link:'should print your relations. Empty output means no header.- Take each target from the header and run
curl -s -o /dev/null -w '%{http_code}\n' https://yoursite.com/llms.txt. Each must print 200. - Fetch the homepage as a bot would, without your browser cookies, since some CDNs vary the headers by user agent.
WordPress. ThinkRank generates the robots.txt, sitemaps, schema and llms.txt that the header points at from one plugin, and it migrates from Rank Math, Yoast, AIOSEO and SEOPress without re-entering settings. The header itself is a server or CDN setting, so add it with one of the paths above. Yoast and Rank Math also generate sitemaps and are the better choice if you already use their other features.
Shopify. StoreSEO generates llms.txt from live products, collections, pages and articles and has an agents.md editor, so the files exist. Whether you can attach a response header depends on your setup, so run the curl check before promising yourself one. A Cloudflare rule in front of a custom domain is the route that does not depend on the theme.
Where AIScan fits, and where it stops
D3 reads the homepage response only. It does not read Link: headers on inner pages, so a per-page describedby pointing at a matching markdown file is invisible to the scan. It also does not judge whether the relations you chose are the right ones, and it cannot see what an agent does after it reads them. Whether any given crawler acts on these headers is not something we measured.
Our own header, listed above, is the worked example. It has four relations and every target returns 200.
Next step
Run a scan and read D3, P1 and C2 together. If D3 fails, paste one line into your server config and scan again. The full list of fixes is on the discoverability checks page, and every step-by-step guide is on /guides. For the wider picture, see six discovery surfaces compared.
Frequently asked questions
What is a Link header for AI agent discovery?
It is an HTTP response header, defined by RFC 8288, that lists URLs related to the page and names each relationship with a rel value. For agent discovery you use it to point at your api-catalog, llms.txt and sitemap so an agent does not have to guess the paths.
Which rel values does AIScan check D3 look for?
AIScan's documentation says D3 inspects the Link response header on your homepage and looks for a rel=api-catalog or rel=describedby entry. It is informational on content sites that have no API catalog.
My curl shows no Link header at all. What do I check first?
Confirm the config was reloaded, then request the exact URL a bot would use, including the trailing slash and the canonical host. In nginx the usual cause is an add_header at a deeper level that cancels the inherited ones, so repeat the Link lines in the location that serves the homepage.
The header shows up in curl but the scan still fails D3. Why?
Your browser or curl may be seeing a different response than the scanner, often because a CDN or cache varies headers by user agent or serves a stale copy. Purge the cache for the homepage, then run curl without cookies and scan again.
Can I send several relations in one Link header?
Yes. RFC 8288 allows one header with comma-separated link values, or several separate Link header lines, and recipients must treat both the same. If another layer already sends a Link header, use append (Apache) or the add operation (Cloudflare) so yours does not replace it.
The scan says my api-catalog relation points at a 404. How do I fix it?
The header only announces the file, so the file must exist. Publish /.well-known/api-catalog per RFC 9727 or remove the relation until you have one, because a header pointing at nothing is worse than no header.
Do I need a Link header if I already have llms.txt and a sitemap?
No file stops working without it, and D3 is informational on content sites. The header removes guesswork for agents that read headers before bodies, and it costs one line of config, so it is worth adding once your target files return 200.
Does the header work on every page or only the homepage?
You can send it on any response, but AIScan D3 reads the homepage only. Per-page relations are valid under RFC 8288 and are not graded by the scan, so verify those by hand with curl -sI on the page URL.
Related guides
Audit Your Own Agent Readiness From a Terminal in Ten Minutes
Paste your URL at aiscan.sitehttps://aiscan.site/ or run npx aiscancli yoursite.com and you get a score in under thirty seconds. Fine for a first read, not enough if you actually want to know what…
The AI-Crawler Statistics Everyone Quotes, and Which You Can Actually Check
Every AIreadiness article, including plenty of ours, opens with a number: 78% of sites have a robots.txt, 28% publish an llms.txt, AI crawlers can't read JavaScript, one crawler is spoofed 17 times…
Publish an API Catalog Agents Can Discover (RFC 9727)
Agents that want to call your API have to guess the URL today. RFC 9727https://www.rfceditor.org/rfc/rfc9727.html gives them a fixed place to look instead: /.wellknown/apicatalog, a small JSON…
The Soft 404 Problem: Why HTTP 200 Breaks Every Agent Discovery Check
Every convention on the agentic web works the same way: an agent asks for a file at a wellknown address, and a 200 means the file is there. /llms.txt, /robots.txt, a .md suffix on a documentation…
