Primafonte

The state of citability on the web

Every month we publish what Primafonte finds across the websites it analyses: how many are ready to be cited by a model, how many are not, and exactly what fails most.

These are aggregate figures from real scans. No individual website or score appears: what is published describes the whole, not anyone in particular.

Sample

8 websites analysed in August 2026, measured with methodology 0.14.0 · provisional thresholds.

The sample for this month is small: 8 websites. The figures are real scans with no estimates involved, but they should be read as what they are — a small count — and not as the state of the sector.

Figures are grouped by methodology version. The criteria have changed over time, so mixing versions would make month-to-month comparison meaningless.

How are the scores distributed?

Scores are not grouped by tens but by what they mean: whether a model can use that website as a source, whether it finds it but cannot cite it reliably, whether it works with specific blind spots, or whether it has no structural blockers. The band matters more than the exact number, because it is what says whether there is anything to fix.

Ready — no structural blockers
3 · 38%
Citable with gaps — it works, with specific blind spots
3 · 38%
Partially readable — they find you, but cannot cite you reliably
2 · 25%
Not citable — models cannot use this site as a source
0 · 0%

What fails most

Share of websites that do not pass each check, calculated only over those that could be measured. Inconclusive results are excluded, as they are in the score.

Question-shaped headings
100% · 8/8
Readable contact details
100% · 8/8
Direct answers after each heading
100% · 8/8
Markdown content negotiation
88% · 7/8
Declared authorship
88% · 7/8
Link discovery headers
88% · 7/8
Data in the content
75% · 6/8
Links to external sources
75% · 6/8
Publication and revision date
75% · 6/8
A declared stance towards the crawlers that cite you
63% · 5/8

What does almost everyone get right?

Most websites pass these five, and seeing them matters as much as seeing the failures: a check everyone passes distinguishes nobody, and if one sits there long enough we have to decide whether it still contributes or is merely holding weight that another would use better.

Time to first byte
100% · 8/8
The crawler is not blocked
100% · 8/8
Render ratio
100% · 7/7
Structured data present and valid
100% · 8/8
Heading hierarchy
100% · 8/8

Which other months have data?

Every month is published separately and with the methodology version it was measured under. Versions are never mixed: the criteria have changed several times, and joining months measured under different rules would give a series that means nothing however tidy it looks.

Where do these figures come from?

The numbers on this page come from real scans run by the engine documented in the methodology. The standards and studies those checks are built on are these, and anyone can check them independently.

  • GEO: Generative Engine Optimization

    The study that measured the effect of citing sources, supplying data and adding quotations: up to 40% more visibility in answers.

  • schema.org

    The vocabulary structured data is validated against, which two of the checks counted here come from.

  • RFC 9309

    The robots.txt standard, which the three crawler checks come from.

  • llms.txt

    The index-file convention for agents, still without a settled specification.

Calibration study

Besides the scans people request, we measured a sample of our own to calibrate the methodology thresholds. These figures come from that sample and are published separately, because we chose the websites ourselves and that changes what the numbers mean.

321 domains from Spain, the United Kingdom, the United States and the rest of Europe, spanning brands, services and mid-sized companies, measured on August 4, 2026 with methodology 0.7.0 · provisional thresholds. Figures are calculated over the 226 that could be measured in full.

51 could not be reached: their firewall blocked the analysis. That is 18% of the attempts that got anywhere. Another 41 did not respond or had no content to analyse.

61

Average score across the sample

How are the sample scores distributed?

The sample was put together by hand, sector by sector, so this distribution describes this set of websites and no other. It is not a snapshot of the web: it is a snapshot of the websites we picked in order to calibrate, which is a rather different thing and worth not confusing.

Ready — no structural blockers
9 · 4%
Citable with gaps — it works, with specific blind spots
124 · 55%
Partially readable — they find you, but cannot cite you reliably
80 · 35%
Not citable — models cannot use this site as a source
13 · 6%

What fails most

Share of websites that do not pass each check, over those that could be measured. Beside it, how many websites that share covers.

A declared stance towards the crawlers that cite you
97% · 215/222
Markdown content negotiation
96% · 217/226
Direct answers after each heading
95% · 194/204
Structured frequently asked questions
94% · 212/226
Entity and business types declared
91% · 206/226
Link discovery headers
85% · 192/226
llms.txt file
82% · 182/222
Paragraphs that stand outside their context
66% · 149/226

What does almost the whole sample get right?

These are the checks almost nobody fails. They are a counterweight to the list above: publishing only what fails would paint a darker picture than the real one, and a methodology that shows only the bad news looks too much like a sales argument.

The crawler is not blocked
100% · 226/226
Time to first byte
95% · 215/226
The crawlers that cite you are not blocked
94% · 209/222
robots.txt reachable and parseable
91% · 202/222
Same content for bots and for people
88% · 172/195
Brand name consistency
85% · 160/188
Identifiable company or author page
80% · 181/226
Clean redirect chain
79% · 179/226

The sample was put together by hand, sector by sector, so it is not representative of the web at large: it describes this set of websites and no other. The list of domains is published in the project repository, and no domain appears here with its score.