AI readiness for news sites and magazines

Your archive is the training set. Decide the terms deliberately, then be citable.

B1B2C3E5D2C2

The moment

Amara, digital editor at a regional title

A reader asks an assistant what happened with a local planning decision this week and gets an answer sourced from an aggregator that rewrote the paper's own reporting.

What the agent needs: The facts, the date, and who reported it.

What actually happens: The article is behind an interstitial, the byline is a script-injected widget, and the site's robots.txt blocks everything indiscriminately after a scraping scare. The aggregator, which blocks nothing, becomes the source.

Which dimensions decide it, and why

AIScan grades five dimensions. These are the ones that carry the outcome for this kind of business.

Bot access

This is the category where the AI position is a strategy, not a checkbox. Content Signals let a publisher permit search and assistant citation while declining training — the all-or-nothing block is what hands attribution to whoever copied you.

Content

Article structured data with headline, datePublished and author is what makes a citation name you. Bylines injected client-side attribute your work to nobody.

Discoverability

A feed and a current sitemap are how a news archive is noticed at the speed news moves. Both are cheap and both are frequently broken on large archives.

The checks that matter here

Fix guides: /docs/checks/bot-access · /docs/checks/content · /docs/platforms/wordpress

Most sites like this run on WordPress read the WordPress playbook for where each fix lives.

Questions we get asked

Should a publisher block AI crawlers outright?
Rarely, and not with one rule. Blocking training crawlers is a licensing position; blocking assistant fetchers removes you from answers that would otherwise cite and link you.
What are Content Signals?
A way to state per-purpose permissions — search, AI input, AI training — inside robots.txt. Check B1 grades whether you have said anything at all.
How do we get cited rather than paraphrased?
Server-rendered article text, Article structured data with a real author, stable URLs and a feed. Citation follows legibility.
Does a paywall have to break this?
No. Serve a substantial, honest preview with structured data marking the paywalled portion, rather than an interstitial that returns nothing.
Our archive is enormous. Where do we start?
robots.txt and the feed first — they are one file each and apply to everything. Then the article template, which fixes every post at once.
How do we monitor for regressions?
Put a scheduled scan on the site so a template change that drops structured data emails you rather than being discovered a quarter later.

Scan your site and see where you actually stand

Free, no account, about twenty seconds. The report names every failing check by ID, shows the evidence we found, and gives the fix for your platform — plus a hand-off prompt you can paste straight into Claude Code or Cursor.

Run a free scan npx aiscan-cli yoursite.com

Prefer to read first? Browse every check we run or the guide library.

Next in Publishing & mediaAI readiness for independent writers and blogs

Also covers: news site, newspaper, magazine, local news, trade publication, online journal, editorial site, content publisher, media company, newsletter archive, paywall, syndication, press site, review site, aggregator, attribution, training opt out.