The state of citability on the web
Every month we publish what Primafonte finds across the websites it analyses: how many are ready to be cited by a model, how many are not, and exactly what fails most.
These are aggregate figures from real scans. No individual website or score appears: what is published describes the whole, not anyone in particular.
Sample
8 websites analysed in August 2026, measured with methodology 0.3.0 · first calibration.
The sample for this month is small: 8 websites. The figures are real scans with no estimates involved, but they should be read as what they are — a small count — and not as the state of the sector.
Figures are grouped by methodology version. The criteria have changed over time, so mixing versions would make month-to-month comparison meaningless.
How the scores are distributed
- Ready — no structural blockers
- 0 · 0%
- Citable with gaps — it works, with specific blind spots
- 7 · 88%
- Partially readable — they find you, but cannot cite you reliably
- 1 · 13%
- Not citable — models cannot use this site as a source
- 0 · 0%
What fails most
Share of websites that do not pass each check, calculated only over those that could be measured. Inconclusive results are excluded, as they are in the score.
- A declared stance towards the crawlers that cite you
- 100% · 8/8
- Entity and business types declared
- 100% · 8/8
- Direct answers after each heading
- 100% · 8/8
- llms.txt file
- 88% · 7/8
- Markdown content negotiation
- 88% · 7/8
- Modification dates in the sitemap
- 80% · 4/5
- Link discovery headers
- 75% · 6/8
- Paragraphs that stand outside their context
- 67% · 4/6
- Title and meta description length
- 63% · 5/8
- Render ratio
- 50% · 4/8
What almost everyone gets right
- The crawler is not blocked
- 100% · 8/8
- jsonld_valid
- 100% · 7/7
- Structured data present and valid
- 87% · 7/8
- robots.txt reachable and parseable
- 87% · 7/8
- The crawlers that cite you are not blocked
- 87% · 7/8
Calibration study
Besides the scans people request, we measured a sample of our own to calibrate the methodology thresholds. These figures come from that sample and are published separately, because we chose the websites ourselves and that changes what the numbers mean.
321 domains from Spain, the United Kingdom, the United States and the rest of Europe, spanning brands, services and mid-sized companies, measured on August 4, 2026 with methodology 0.7.0 · first calibration. Figures are calculated over the 226 that could be measured in full.
51 could not be reached: their firewall blocked the analysis. That is 18% of the attempts that got anywhere. Another 41 did not respond or had no content to analyse.
61
Average score across the sample
How the scores break down
- Ready — no structural blockers
- 9 · 4%
- Citable with gaps — it works, with specific blind spots
- 124 · 55%
- Partially readable — they find you, but cannot cite you reliably
- 80 · 35%
- Not citable — models cannot use this site as a source
- 13 · 6%
What fails most
Share of websites that do not pass each check, over those that could be measured. Beside it, how many websites that share covers.
- A declared stance towards the crawlers that cite you
- 97% · 215/222
- Markdown content negotiation
- 96% · 217/226
- Direct answers after each heading
- 95% · 194/204
- Structured frequently asked questions
- 94% · 212/226
- Entity and business types declared
- 91% · 206/226
- Link discovery headers
- 85% · 192/226
- llms.txt file
- 82% · 182/222
- Paragraphs that stand outside their context
- 66% · 149/226
What almost everyone gets right
- The crawler is not blocked
- 100% · 226/226
- Time to first byte
- 95% · 215/226
- The crawlers that cite you are not blocked
- 94% · 209/222
- robots.txt reachable and parseable
- 91% · 202/222
- Same content for bots and for people
- 88% · 172/195
- Brand name consistency
- 85% · 160/188
- Identifiable company or author page
- 80% · 181/226
- Clean redirect chain
- 79% · 179/226
The sample was put together by hand, sector by sector, so it is not representative of the web at large: it describes this set of websites and no other. The list of domains is published in the project repository, and no domain appears here with its score.