AI readiness for news sites and magazines
Your archive is the training set. Decide the terms deliberately, then be citable.
The moment
Amara, digital editor at a regional title
A reader asks an assistant what happened with a local planning decision this week and gets an answer sourced from an aggregator that rewrote the paper's own reporting.
What the agent needs: The facts, the date, and who reported it.
What actually happens: The article is behind an interstitial, the byline is a script-injected widget, and the site's robots.txt blocks everything indiscriminately after a scraping scare. The aggregator, which blocks nothing, becomes the source.
Which dimensions decide it, and why
AIScan grades five dimensions. These are the ones that carry the outcome for this kind of business.
This is the category where the AI position is a strategy, not a checkbox. Content Signals let a publisher permit search and assistant citation while declining training — the all-or-nothing block is what hands attribution to whoever copied you.
Article structured data with headline, datePublished and author is what makes a citation name you. Bylines injected client-side attribute your work to nobody.
A feed and a current sitemap are how a news archive is noticed at the speed news moves. Both are cheap and both are frequently broken on large archives.
The checks that matter here
- B1Content Signals in robots.txt
- B2Explicit AI bot rules
- C3Structured HTML (title, meta, JSON-LD, single H1)
- E5Content feed (RSS / Atom / JSON Feed)
- D2XML sitemap
- C2/llms.txt
Fix guides: /docs/checks/bot-access · /docs/checks/content · /docs/platforms/wordpress
Most sites like this run on WordPress — read the WordPress playbook for where each fix lives.
Questions we get asked
Should a publisher block AI crawlers outright?
What are Content Signals?
How do we get cited rather than paraphrased?
Does a paywall have to break this?
Our archive is enormous. Where do we start?
How do we monitor for regressions?
Scan your site and see where you actually stand
Free, no account, about twenty seconds. The report names every failing check by ID, shows the evidence we found, and gives the fix for your platform — plus a hand-off prompt you can paste straight into Claude Code or Cursor.
npx aiscan-cli yoursite.comPrefer to read first? Browse every check we run or the guide library.
Also covers: news site, newspaper, magazine, local news, trade publication, online journal, editorial site, content publisher, media company, newsletter archive, paywall, syndication, press site, review site, aggregator, attribution, training opt out.