The state of citability on the web
Every month we publish what Primafonte finds across the websites it analyses: how many are ready to be cited by a model, how many are not, and exactly what fails most.
These are aggregate figures from real scans. No individual website or score appears: what is published describes the whole, not anyone in particular.
Sample
8 websites analysed in August 2026, measured with methodology 0.14.0 · provisional thresholds.
The sample for this month is small: 8 websites. The figures are real scans with no estimates involved, but they should be read as what they are — a small count — and not as the state of the sector.
Figures are grouped by methodology version. The criteria have changed over time, so mixing versions would make month-to-month comparison meaningless.
How are the scores distributed?
Scores are not grouped by tens but by what they mean: whether a model can use that website as a source, whether it finds it but cannot cite it reliably, whether it works with specific blind spots, or whether it has no structural blockers. The band matters more than the exact number, because it is what says whether there is anything to fix.
- Ready — no structural blockers
- 3 · 38%
- Citable with gaps — it works, with specific blind spots
- 3 · 38%
- Partially readable — they find you, but cannot cite you reliably
- 2 · 25%
- Not citable — models cannot use this site as a source
- 0 · 0%
What fails most
Share of websites that do not pass each check, calculated only over those that could be measured. Inconclusive results are excluded, as they are in the score.
- Question-shaped headings
- 100% · 8/8
- Readable contact details
- 100% · 8/8
- Direct answers after each heading
- 100% · 8/8
- Markdown content negotiation
- 88% · 7/8
- Declared authorship
- 88% · 7/8
- Link discovery headers
- 88% · 7/8
- Data in the content
- 75% · 6/8
- Links to external sources
- 75% · 6/8
- Publication and revision date
- 75% · 6/8
- A declared stance towards the crawlers that cite you
- 63% · 5/8
What does almost everyone get right?
Most websites pass these five, and seeing them matters as much as seeing the failures: a check everyone passes distinguishes nobody, and if one sits there long enough we have to decide whether it still contributes or is merely holding weight that another would use better.
- Time to first byte
- 100% · 8/8
- The crawler is not blocked
- 100% · 8/8
- Render ratio
- 100% · 7/7
- Structured data present and valid
- 100% · 8/8
- Heading hierarchy
- 100% · 8/8
Which other months have data?
Every month is published separately and with the methodology version it was measured under. Versions are never mixed: the criteria have changed several times, and joining months measured under different rules would give a series that means nothing however tidy it looks.
Where do these figures come from?
The numbers on this page come from real scans run by the engine documented in the methodology. The standards and studies those checks are built on are these, and anyone can check them independently.
- GEO: Generative Engine Optimization
The study that measured the effect of citing sources, supplying data and adding quotations: up to 40% more visibility in answers.
- schema.org
The vocabulary structured data is validated against, which two of the checks counted here come from.
- RFC 9309
The robots.txt standard, which the three crawler checks come from.
- llms.txt
The index-file convention for agents, still without a settled specification.
Calibration study
Besides the scans people request, we measured a sample of our own to calibrate the methodology thresholds. These figures come from that sample and are published separately, because we chose the websites ourselves and that changes what the numbers mean.
321 domains from Spain, the United Kingdom, the United States and the rest of Europe, spanning brands, services and mid-sized companies, measured on August 4, 2026 with methodology 0.7.0 · provisional thresholds. Figures are calculated over the 226 that could be measured in full.
51 could not be reached: their firewall blocked the analysis. That is 18% of the attempts that got anywhere. Another 41 did not respond or had no content to analyse.
61
Average score across the sample
How are the sample scores distributed?
The sample was put together by hand, sector by sector, so this distribution describes this set of websites and no other. It is not a snapshot of the web: it is a snapshot of the websites we picked in order to calibrate, which is a rather different thing and worth not confusing.
- Ready — no structural blockers
- 9 · 4%
- Citable with gaps — it works, with specific blind spots
- 124 · 55%
- Partially readable — they find you, but cannot cite you reliably
- 80 · 35%
- Not citable — models cannot use this site as a source
- 13 · 6%
What fails most
Share of websites that do not pass each check, over those that could be measured. Beside it, how many websites that share covers.
- A declared stance towards the crawlers that cite you
- 97% · 215/222
- Markdown content negotiation
- 96% · 217/226
- Direct answers after each heading
- 95% · 194/204
- Structured frequently asked questions
- 94% · 212/226
- Entity and business types declared
- 91% · 206/226
- Link discovery headers
- 85% · 192/226
- llms.txt file
- 82% · 182/222
- Paragraphs that stand outside their context
- 66% · 149/226
What does almost the whole sample get right?
These are the checks almost nobody fails. They are a counterweight to the list above: publishing only what fails would paint a darker picture than the real one, and a methodology that shows only the bad news looks too much like a sales argument.
- The crawler is not blocked
- 100% · 226/226
- Time to first byte
- 95% · 215/226
- The crawlers that cite you are not blocked
- 94% · 209/222
- robots.txt reachable and parseable
- 91% · 202/222
- Same content for bots and for people
- 88% · 172/195
- Brand name consistency
- 85% · 160/188
- Identifiable company or author page
- 80% · 181/226
- Clean redirect chain
- 79% · 179/226
The sample was put together by hand, sector by sector, so it is not representative of the web at large: it describes this set of websites and no other. The list of domains is published in the project repository, and no domain appears here with its score.