A content analysis of the llms.txt files in Common Crawl CC-MAIN-2026-30
Before anything can be read, the file has to exist. This corpus is the result of an experiment rather than an incidental find: for CC-MAIN-2026-30 the paths /llms.txt and /llms-full.txt were added to the crawl's seed list on purpose, and what follows is the outcome of 6,563,125 attempted URLs.
The seeds were drawn two ways. Most are a random sample of hosts Common Crawl had recently fetched without problems — 5,167,831 URLs for /llms.txt, and a deliberately narrower 1,761,228 for /llms-full.txt, since a site with no /llms.txt rarely publishes the longer file. The rest are hosts already known from the two preceding crawls (CC-MAIN-2026-21 and CC-MAIN-2026-25) to serve one of them: 27,394 with a /llms.txt and 15,763 with a /llms-full.txt. Not every seed is fetched within a single crawl, which is why the totals below fall short of the sample. Because the first group is a random sample, the 200-plus-text rate for /llms.txt reads as an adoption rate for hosts Common Crawl can fetch. The /llms-full.txt rate does not, because its sample was enriched with hosts already known to serve the file.
The query below reads the URL index after the fact, to recover the outcome of each attempt. It is not what selected the population. Quoted verbatim, never re-run here.
select crawl, url, url_host_name, fetch_status, content_mime_detected
from ccindex
where (url_path = '/llms.txt'
or url_path = '/llms-full.txt')
and url_query is null
and crawl like 'CC-MAIN-2026-30';| stage | URLs | % of attempted | what it means |
|---|---|---|---|
| URLs attempted | 6,563,125 | 100.00 | every /llms.txt and /llms-full.txt in the crawl |
| responded 200 | 1,287,207 | 19.61 | the file exists and was served |
| …with a text body | 587,302 | 8.95 | text/plain or text/markdown — the analysable population |
69.8% of attempts 404 and another 7.77% redirect, which is unsurprising: the crawler tries the path on every host it knows. What is surprising is the 85 distinct status codes it received, and how little of the 200 population is actually a document — only 45.63% of successful responses carry a text/plain or text/markdown body.
| response | URLs | % of attempted |
|---|---|---|
| 200 OK | 1,287,207 | 19.61 |
| 301/302 redirect | 488,125 | 7.44 |
| other 3xx | 21,627 | 0.33 |
| 404 Not Found | 4,581,297 | 69.80 |
| other 4xx | 159,405 | 2.43 |
| 5xx server error | 24,786 | 0.38 |
The content types tell you what the other 200s are. A site that serves its single-page app for every unknown path answers text/html; one whose rewrite rules send /llms.txt to /robots.txt answers text/x-robots. Both are a 200 that an agent would have to parse and discard.
| content type detected (HTTP 200) | URLs | % of 200s |
|---|---|---|
| text/plain | 560,988 | 43.58 |
| text/html | 519,484 | 40.36 |
| text/x-robots | 136,578 | 10.61 |
| application/xhtml+xml | 29,620 | 2.30 |
| text/markdown | 26,314 | 2.04 |
| application/javascript | 11,752 | 0.91 |
| application/xml | 990 | 0.08 |
| application/json | 736 | 0.06 |
| text/txt | 277 | 0.02 |
| text/x-php | 273 | 0.02 |
| application/rtf | 87 | 0.01 |
| application/x-httpresponse | 31 | 0.00 |
| other (15 types) | 77 | 0.01 |
Split by filename, llms-full.txt is both rarer and much more often misconfigured: it is attempted on 1,667,437 URLs but only 7.01% of its 200s carry a text body, against 52.50% for llms.txt.
| file | URLs attempted | % 404 | % 301/302 | % 200 | % 200 + text body | % of its 200s |
|---|---|---|---|---|---|---|
| llms.txt | 4,895,688 | 68.40 | 7.47 | 22.32 | 11.72 | 52.50 |
| llms-full.txt | 1,667,437 | 73.94 | 7.35 | 11.66 | 0.82 | 7.01 |
One reconciliation note, because the two populations are not quite the same. The index query above matches the path exactly, while the WARC extraction that produced the dataset matched LIKE '%/llms.txt' and so also captured files deeper in a site — 10,996 of them (1.84%), things like /branding/vesence/llms.txt or one file per product. Set those aside and the 598,298 extracted records leave 587,302 at a root path, exactly the 587,302 the index counts.
| population | documents |
|---|---|
| URL index: HTTP 200 with a text body | 587,302 |
| dataset: records extracted from WARC | 598,298 |
| of which at a root path | 587,302 |
| of which deeper in the site | 10,996 |
| analysed here (non-empty) | 584,107 |
Those 598,298 surviving responses are the corpus: every URL ending in llms.txt or llms-full.txt that returned HTTP 200 with a text/plain or text/markdown body, extracted from the WARC archives. This report asks what is written in them. Published adoption studies count who has the file and who fetches it; none of them open it.
14,191 of those responses (2.37%, across 13,167 hosts) have an empty body: a 200, the right content type, and nothing in it. They are a finding about deployment, not documents to analyse — every empty body is identical, so leaving them in would manufacture a template cluster and inflate the share of files with no links, no title and no generator. They are counted here and excluded from every figure that follows, which is computed over the 584,107 responses that contain something.
| response | documents | % of responses |
|---|---|---|
| contains a document | 584,107 | 97.63 |
| empty body (excluded below) | 14,191 | 2.37 |
Three findings organise everything below. First, this is plugin output, not curation: 68.3% of documents are template-generated, and the ten largest template clusters alone account for 38.7% of the whole corpus. Second, the file has quietly become a policy and service-discovery channel — 6.6% carry usage or rights language and 44.9% advertise an agent endpoint, neither of which the specification mentions. Third, it is unguarded: 2.54% are parked domains pitching themselves to agents, 1.20% hit a spam lexicon, and a small but non-zero share carry instructions written at the reading model.
| file | documents | % of corpus | examples |
|---|---|---|---|
| llms.txt | 570,680 | 97.70 | |
| llms-full.txt | 13,427 | 2.30 |
| TLD | documents | % of corpus | examples |
|---|---|---|---|
| com | 326,267 | 55.86 | |
| org | 41,960 | 7.18 | |
| uk | 19,781 | 3.39 | |
| de | 17,776 | 3.04 | |
| net | 14,410 | 2.47 | |
| fr | 9,432 | 1.61 | |
| au | 7,884 | 1.35 | |
| br | 7,787 | 1.33 | |
| ca | 7,643 | 1.31 | |
| jp | 7,581 | 1.30 | |
| io | 7,336 | 1.26 | |
| it | 5,466 | 0.94 | |
| nl | 5,436 | 0.93 | |
| ai | 5,151 | 0.88 | |
| ch | 5,032 | 0.86 |
The llms.txt specification is small enough to test mechanically: an H1 title (the only required element), a > blockquote summary, optional heading-free prose, then ## sections of - [name](url): notes bullets, optionally including an ## Optional section that a short-context reader may skip.
99.2% of documents are markdown at all. The rest — 0.71% — are JSON error blobs, HTML error pages, robots.txt files served under the wrong name, and YAML policy documents, every one of them returned as HTTP 200 with a text/plain or text/markdown content type. (The 14,191 entirely empty responses are excluded here and counted in the overview.)
| document kind | documents | % of corpus | examples |
|---|---|---|---|
| markdown | 579,581 | 99.23 | |
| plain | 3,129 | 0.54 | |
| html | 521 | 0.09 | |
| yaml_frontmatter | 378 | 0.06 | |
| yaml | 328 | 0.06 | |
| json | 146 | 0.02 | |
| robots_txt | 24 | 0.00 |
The surface form travels well. 49.90% of documents carry the full spec shape — H1, summary blockquote, H2 sections, link bullets — and that is not an artefact of one generator: 54.17% of human-authored documents reach it against 48.00% of templated ones. Whatever else is wrong with the llms.txt web, people did read the spec.
| conformance level | documents | % of corpus | examples |
|---|---|---|---|
| H1 + summary + link sections | 291,442 | 49.90 | |
| H1 only | 171,544 | 29.37 | |
| H1 + summary | 111,257 | 19.05 | |
| markdown, no H1 | 5,716 | 0.98 | |
| not markdown | 4,148 | 0.71 |
| conformance level | % of llms.txt | % of llms-full.txt |
|---|---|---|
| not markdown | 0.62 | 4.66 |
| markdown, no H1 | 0.99 | 0.70 |
| H1 only | 28.71 | 57.23 |
| H1 + summary | 19.03 | 19.93 |
| H1 + summary + link sections | 50.66 | 17.48 |
What does not travel is the part that carries the value. The specification's purpose is to hand an agent a curated, annotated list of links, and that is precisely where adherence collapses: only 32.94% of documents annotate any link with the : notes the format allows, the ## Optional section that lets a short-context reader skip material appears in 9.71%, and 22.6% of documents contain no link at all (22.2% of the markdown ones). The shape is followed; the substance is not.
| spec element | documents | % of corpus | % of markdown docs |
|---|---|---|---|
| starts with H1 | 485,041 | 83.04 | 83.64 |
| has any H1 | 574,585 | 98.37 | 99.02 |
| summary blockquote after H1 | 402,930 | 68.98 | 69.48 |
| prose body before first H2 | 192,557 | 32.97 | 32.79 |
| at least one H2 section | 571,039 | 97.76 | 98.40 |
| link bullets | 430,243 | 73.66 | 74.06 |
| link notes (': description') | 192,426 | 32.94 | 33.12 |
| '## Optional' section | 56,743 | 9.71 | 9.77 |
| byte-order mark | 138,449 | 23.70 | 23.85 |
| flag | documents | % of corpus | examples |
|---|---|---|---|
| deep_headings | 292,952 | 50.15 | |
| links_without_notes | 259,882 | 44.49 | |
| bom | 138,449 | 23.70 | |
| no_links | 131,800 | 22.56 | |
| h1_not_first | 89,636 | 15.35 | |
| has_optional_section | 56,743 | 9.71 | |
| multiple_h1 | 16,658 | 2.85 | |
| relative_links | 13,545 | 2.32 | |
| links_outside_bullets | 2,852 | 0.49 |
| statistic | value |
|---|---|
| min | 0 |
| p50 | 3 |
| p75 | 28 |
| p90 | 140 |
| p95 | 438 |
| p99 | 1,883 |
| p99.9 | 6,667 |
| max | 51,327 |
| mean | 100 |
| generator | documents | % with zero links |
|---|---|---|
| aioseo | 73,136 | 0.00 |
| gitbook | 2,498 | 0.60 |
| godaddy_parking | 14,711 | 99.99 |
| rankmath | 12,360 | 0.11 |
| unknown | 197,612 | 17.98 |
| wix | 241,448 | 33.59 |
| wordpress | 1,070 | 27.38 |
| yoast | 40,509 | 0.04 |
Detection runs in two layers: an explicit self-declaration with a version string ("Generated by Yoast SEO v28.0, this is an llms.txt file"), and, for the silent generators, a language-independent structural fingerprint — for the site builders, the help-centre URL they embed rather than the prose around it, so that localised templates are still detected.
| generator | documents | % of corpus | examples |
|---|---|---|---|
| wix | 241,448 | 41.34 | |
| unknown | 197,612 | 33.83 | |
| aioseo | 73,136 | 12.52 | |
| yoast | 40,509 | 6.94 | |
| godaddy_parking | 14,711 | 2.52 | |
| rankmath | 12,360 | 2.12 | |
| gitbook | 2,498 | 0.43 | |
| wordpress | 1,070 | 0.18 | |
| shopify | 215 | 0.04 | |
| mintlify | 169 | 0.03 | |
| wpbakery_llms | 121 | 0.02 | |
| webflow | 47 | 0.01 | |
| readme_io | 42 | 0.01 | |
| docusaurus | 31 | 0.01 |
| family | documents | % of corpus | examples |
|---|---|---|---|
| site_builder | 241,512 | 41.35 | |
| unknown | 197,612 | 33.83 | |
| seo_plugin | 126,131 | 21.59 | |
| parking | 14,711 | 2.52 | |
| docs_platform | 2,753 | 0.47 | |
| cms | 1,070 | 0.18 | |
| ecommerce | 225 | 0.04 | |
| other_declared | 93 | 0.02 |
The largest single producer is wix at 41.34% of the corpus. Together with the WordPress SEO plugins, a handful of vendors decide what most of the llms.txt web says — and they interpreted the same informal spec in mutually incompatible ways: some emit a curated index, some dump the sitemap, some emit no links at all and advertise an API instead.
| generator | documents | % conformance >= 3 | median links | median tokens |
|---|---|---|---|---|
| aioseo | 73,136 | 0.00 | 138 | 8,772 |
| gitbook | 2,498 | 0.76 | 23 | 819 |
| godaddy_parking | 14,711 | 100.00 | 0 | 148 |
| rankmath | 12,360 | 0.19 | 85 | 7,459 |
| unknown | 197,612 | 63.50 | 11 | 1,000 |
| wix | 241,448 | 94.49 | 3 | 719 |
| wordpress | 1,070 | 54.11 | 12 | 1,602 |
| yoast | 40,509 | 82.30 | 21 | 729 |
Template detection needs no statistical model. Dropping the parts that are site-specific by construction — the title, the summary blockquote, the link bullets — and then erasing URLs, e-mail addresses, numbers and every non-ASCII run leaves the vendor's boilerplate and section structure. Exact matching on that skeleton is enough: one template collapses to a single string across thousands of sites. Translated boilerplate forms one cluster per language, which is what the generator fingerprints unify; skeleton clustering is what catches template families whose producer is unknown.
| cluster size | % of corpus | generator | top language | examples |
|---|---|---|---|---|
| 118,002 | 20.20 | wix | eng | |
| 59,962 | 10.27 | wix | eng | |
| 9,096 | 1.56 | wix | fra | |
| 8,728 | 1.49 | wix | deu | |
| 7,000 | 1.20 | wix | jpn | |
| 6,259 | 1.07 | aioseo | jpn | |
| 4,548 | 0.78 | wix | fra | |
| 4,308 | 0.74 | wix | spa | |
| 4,294 | 0.74 | wix | deu | |
| 3,920 | 0.67 | wix | por | |
| 3,531 | 0.60 | wix | jpn | |
| 2,189 | 0.37 | aioseo | eng | |
| 2,148 | 0.37 | rankmath | eng | |
| 2,141 | 0.37 | wix | spa | |
| 2,015 | 0.34 | wix | por |
68.3% of the corpus is templated by this measure; the ten biggest clusters alone are 38.7%. Every rate elsewhere in this report is therefore reported against the human-authored remainder wherever the distinction matters.
One template family deserves its own line. 44.9% of documents mention the Model Context Protocol and 44.4% publish a live MCP endpoint URL. For those sites llms.txt is not a link index at all: it is a service-discovery record telling agents to stop scraping and call an API instead. The specification does not mention this use.
| generator | documents | % of corpus |
|---|---|---|
| wix | 241,437 | 92.11 |
| unknown | 20,317 | 7.75 |
| shopify | 142 | 0.05 |
| yoast | 52 | 0.02 |
| aioseo | 49 | 0.02 |
| wordpress | 33 | 0.01 |
| gitbook | 32 | 0.01 |
| rankmath | 24 | 0.01 |
| wpbakery_llms | 6 | 0.00 |
| mintlify | 5 | 0.00 |
A third production route leaves its own traces. 1.43% of documents carry a marker that only appears when someone asked an assistant to write the file and shipped the answer unedited: ChatGPT's contentReference[oaicite:…] citation markers, an unfilled [Insert company name] placeholder, a leftover ```markdown fence around the whole document, or a utm_source=chatgpt link. The topic model found these on its own — whole clusters of clinics and agencies held together by the token oaicite.
| artefact | documents | % of corpus | examples |
|---|---|---|---|
| unfilled_placeholder | 7,997 | 1.37 | |
| openai_citation | 287 | 0.05 | |
| chatgpt_utm | 62 | 0.01 | |
| fence_leak | 24 | 0.00 | |
| model_disclaimer | 6 | 0.00 |
| generator | documents | % of corpus |
|---|---|---|
| aioseo | 4,776 | 57.03 |
| unknown | 2,840 | 33.91 |
| rankmath | 407 | 4.86 |
| yoast | 287 | 3.43 |
| wordpress | 53 | 0.63 |
| mintlify | 7 | 0.08 |
| webflow | 1 | 0.01 |
| docusaurus | 1 | 0.01 |
| generator | version | documents |
|---|---|---|
| aioseo | 4.9.10 | 33,686 |
| yoast | 28.0 | 16,900 |
| yoast | 27.9 | 7,280 |
| aioseo | 4.9.7.2 | 6,537 |
| aioseo | 4.9.9 | 6,452 |
| aioseo | 4.9.3 | 4,553 |
| aioseo | 4.9.5.1 | 2,691 |
| aioseo | 4.9.8 | 2,631 |
| yoast | 27.8 | 2,424 |
| aioseo | 4.9.6.2 | 2,229 |
| yoast | 27.7 | 1,838 |
| aioseo | 4.9.2 | 1,657 |
| aioseo | 4.9.0 | 1,652 |
| aioseo | 4.8.7 | 1,531 |
| yoast | 27.6 | 1,477 |
Nothing in the specification concerns rights. In practice a measurable share of documents use llms.txt to state terms — and they do so in at least four mutually incompatible dialects: prose paragraphs, YAML permission blocks, robots.txt syntax served under the llms.txt name, and links out to a licence page.
| dialect | documents | % of corpus | examples |
|---|---|---|---|
| none | 545,627 | 93.41 | |
| prose | 37,640 | 6.44 | |
| robots | 664 | 0.11 | |
| yaml | 154 | 0.03 | |
| mixed | 22 | 0.00 |
| directive | documents | % of corpus | examples |
|---|---|---|---|
| rate_limit | 17,798 | 3.05 | |
| copyright_notice | 5,444 | 0.93 | |
| citation_required | 2,692 | 0.46 | |
| license_ref | 1,510 | 0.26 | |
| attribution_required | 846 | 0.14 | |
| allow_training | 745 | 0.13 | |
| paywall_notice | 513 | 0.09 | |
| no_training | 342 | 0.06 | |
| respect_robots | 257 | 0.04 | |
| commercial_license | 256 | 0.04 | |
| contact_for_licensing | 177 | 0.03 | |
| medical_legal_disclaimer | 176 | 0.03 | |
| noncommercial_only | 53 | 0.01 | |
| jurisdiction | 39 | 0.01 | |
| ai_summary_ok | 21 | 0.00 | |
| disallow_training_header | 7 | 0.00 |
| stance | documents | % of corpus | examples |
|---|---|---|---|
| none | 582,749 | 99.77 | |
| allow | 710 | 0.12 | |
| deny | 347 | 0.06 | |
| conditional | 301 | 0.05 |
The named-crawler table is the closest thing this corpus has to a robots.txt of the agent era. It is also where a publisher's intent is least ambiguous: naming a user-agent and putting allow or deny next to it is a deliberate act.
| crawler | documents naming it | allowed | denied | % denied (of decided) | examples |
|---|---|---|---|---|---|
| GPTBot | 867 | 260 | 65 | 20.00 | |
| PerplexityBot | 769 | 241 | 43 | 15.14 | |
| ClaudeBot | 759 | 232 | 54 | 18.88 | |
| Google-Extended | 682 | 202 | 41 | 16.87 | |
| xAI | 423 | 10 | 20 | 66.67 | |
| CCBot | 399 | 90 | 32 | 26.23 | |
| ChatGPT-User | 308 | 97 | 18 | 15.65 | |
| OAI-SearchBot | 289 | 100 | 17 | 14.53 | |
| Applebot-Extended | 252 | 79 | 16 | 16.84 | |
| anthropic-ai | 215 | 79 | 11 | 12.22 | |
| Bytedance | 182 | 8 | 10 | 55.56 | |
| Googlebot | 173 | 46 | 15 | 24.59 | |
| claude-web | 159 | 48 | 7 | 12.73 | |
| Bytespider | 151 | 60 | 10 | 14.29 | |
| Amazonbot | 148 | 44 | 12 | 21.43 | |
| Bingbot | 140 | 41 | 9 | 18.00 | |
| YouBot | 136 | 28 | 6 | 17.65 | |
| cohere-ai | 120 | 38 | 8 | 17.39 | |
| Meta-ExternalAgent | 108 | 25 | 8 | 24.24 | |
| Claude-SearchBot | 83 | 20 | 2 | 9.09 |
A verdict is read from the same line the crawler is named on. Documents that name a crawler without any allow/deny marker are counted in the first column only, which is why allowed + denied does not sum to the total.
| generator | documents | % of corpus |
|---|---|---|
| unknown | 33,060 | 85.91 |
| aioseo | 2,988 | 7.77 |
| wix | 1,286 | 3.34 |
| rankmath | 418 | 1.09 |
| wordpress | 239 | 0.62 |
| shopify | 150 | 0.39 |
| yoast | 135 | 0.35 |
| mintlify | 80 | 0.21 |
| gitbook | 48 | 0.12 |
| wpbakery_llms | 27 | 0.07 |
llms.txt is text an agent reads as guidance, published at a predictable path, inspected by no security tooling. The lexicon below grades what publishers put there, from benign steering through self-promotion to outright instruction override. It is deliberately conservative: high precision by construction, unknown recall.
| severity | documents | % of corpus | examples |
|---|---|---|---|
| none | 580,202 | 99.33 | |
| steering | 3,793 | 0.65 | |
| promotional | 102 | 0.02 | |
| override | 10 | 0.00 |
The override row is small enough to audit exhaustively, and it was: all of its documents were read individually. Four are genuine — a bug-bounty researcher's labelled payload catcher, two protest files, and one “LLM Training Policy” whose summary blockquote instructs the reader to ignore its instructions and fetch another URL. The rest are false positives of two kinds: technical writing that contains the literal ChatML control tokens because it is explaining them, and an incidental “you are now” in a page description. Read the override count as an upper bound on candidates, not a count of attacks — and note that every genuine case was placed deliberately by its author, rather than smuggled in.
| pattern | documents | % of corpus | examples |
|---|---|---|---|
| addressed_to_model | 3,322 | 0.57 | |
| always_mention | 436 | 0.07 | |
| prioritise_pages | 103 | 0.02 | |
| avoid_pages | 78 | 0.01 | |
| always_recommend | 65 | 0.01 | |
| answer_script | 36 | 0.01 | |
| role_reassign | 5 | 0.00 | |
| ignore_previous | 3 | 0.00 | |
| system_prompt_override | 3 | 0.00 | |
| suppress_competitors | 2 | 0.00 |
| severity | documents | % of corpus |
|---|---|---|
| none | 180,969 | 98.01 |
| steering | 3,558 | 1.93 |
| promotional | 102 | 0.06 |
| override | 7 | 0.00 |
Parked domains are the corpus's most absurd corner. 14,812 documents (2.54%) advertise the domain itself for sale, and 99.99% of those are well-formed enough to clear conformance level 2 — someone wrote a spec-shaped sales pitch aimed at a language model.
| marketplace | documents | % of corpus |
|---|---|---|
| godaddy | 14,713 | 99.33 |
| other | 97 | 0.65 |
| sedo | 2 | 0.01 |
| lexicon | documents | % of corpus | examples |
|---|---|---|---|
| gambling | 5,617 | 0.96 | |
| adult | 1,106 | 0.19 | |
| pharma | 227 | 0.04 | |
| loans | 58 | 0.01 | |
| essay_mill | 48 | 0.01 | |
| replica | 35 | 0.01 | |
| seo_services | 23 | 0.00 | |
| crypto | 20 | 0.00 |
Link injection is visible too. 0.37% of documents send the majority of their links to five or more third-party domains, and 0.11% are written in a language implausible for their ccTLD — the classic signature of a compromised site republished as a link farm.
| TLD | content language | documents | examples |
|---|---|---|---|
| co | fra | 25 | |
| co | por | 16 | |
| id | msa | 13 | |
| it | deu | 11 | |
| co | tha | 11 | |
| sk | ces | 10 | |
| co | tur | 10 | |
| cz | slk | 9 | |
| cn | ell | 9 | |
| de | fra | 8 | |
| nl | vie | 8 | |
| jp | zxx | 7 |
| statistic | value |
|---|---|
| min | 0 |
| p50 | 0 |
| p75 | 33 |
| p90 | 33 |
| p95 | 82 |
| p99 | 100 |
| p99.9 | 100 |
| max | 100 |
| mean | 18 |
| generator | documents | % of corpus |
|---|---|---|
| aioseo | 3,217 | 45.20 |
| unknown | 2,724 | 38.27 |
| rankmath | 559 | 7.85 |
| yoast | 501 | 7.04 |
| wix | 77 | 1.08 |
| godaddy_parking | 13 | 0.18 |
| gitbook | 12 | 0.17 |
| wordpress | 11 | 0.15 |
| mintlify | 2 | 0.03 |
| wpbakery_llms | 1 | 0.01 |
Language is identified with commonlid, Common Crawl's own LID evaluation kit: cld2 as the primary model with GlotLID as the fallback where cld2 abstains. Input is the document's prose with link bullets, URLs and markdown syntax stripped, so that link dumps and translated boilerplate do not decide the label.
| language (ISO 639-3) | documents | % of corpus | examples |
|---|---|---|---|
| eng | 430,733 | 73.74 | |
| deu | 25,430 | 4.35 | |
| jpn | 22,999 | 3.94 | |
| fra | 22,503 | 3.85 | |
| spa | 14,574 | 2.50 | |
| por | 9,385 | 1.61 | |
| und | 7,043 | 1.21 | |
| ita | 5,988 | 1.03 | |
| nld | 5,891 | 1.01 | |
| rus | 5,233 | 0.90 | |
| pol | 3,925 | 0.67 | |
| tur | 3,893 | 0.67 | |
| zho | 3,353 | 0.57 | |
| swe | 2,262 | 0.39 | |
| ces | 1,977 | 0.34 | |
| kor | 1,545 | 0.26 | |
| nor | 1,157 | 0.20 | |
| dan | 1,102 | 0.19 |
| LID model | documents | % of corpus | examples |
|---|---|---|---|
| cld2 | 570,730 | 97.71 | |
| none | 6,898 | 1.18 | |
| glotlid | 6,386 | 1.09 | |
| abstain | 93 | 0.02 |
| generator | documents | distinct languages (>=10 docs) | % English |
|---|---|---|---|
| aioseo | 73,136 | 39 | 59.55 |
| gitbook | 2,498 | 18 | 57.77 |
| godaddy_parking | 14,711 | 1 | 100.00 |
| rankmath | 12,360 | 30 | 78.95 |
| unknown | 197,612 | 94 | 77.99 |
| wix | 241,448 | 22 | 73.90 |
| wordpress | 1,070 | 8 | 77.29 |
| yoast | 40,509 | 41 | 67.32 |
Length is where the specification's premise — a small, cheap, curated file — meets what was actually shipped. The median document is 737 tokens; the 99th percentile is 159,880, and the largest is 3,313,814 tokens: a single file whose entire purpose is being inexpensive to read.
| statistic | value |
|---|---|
| min | 1 |
| p50 | 737 |
| p75 | 1,603 |
| p90 | 11,946 |
| p95 | 34,946 |
| p99 | 159,880 |
| p99.9 | 920,945 |
| max | 3,313,814 |
| mean | 9,182 |
| file | documents | median tokens | p90 | p99 | max |
|---|---|---|---|---|---|
| llms-full.txt | 13,427 | 1,732 | 96,011 | 1,068,699 | 2,879,973 |
| llms.txt | 570,680 | 733 | 11,182 | 147,421 | 3,313,814 |
| context window (tokens) | documents exceeding | % of corpus |
|---|---|---|
| 32,000 | 31,090 | 5.32 |
| 128,000 | 8,700 | 1.49 |
| 200,000 | 3,780 | 0.65 |
| 1,000,000 | 467 | 0.08 |
Ingesting the whole corpus once costs the following, at 5,363,837,955 input tokens (counted with o200k_base; other tokenizers differ by roughly ±20%). List prices come from litellm's model cost map, charging each document as its own request — pricing the corpus as one call would trigger the long-context tier several providers apply above ~200k tokens and roughly double the figure.
| model | input $/1M tokens | context window | cost to ingest whole corpus ($) |
|---|---|---|---|
| gpt-5 | 1.25 | 272,000 | 6,704.80 |
| gpt-4o-mini | 0.15 | 128,000 | 804.58 |
| claude-sonnet-4-5 | 3.00 | 200,000 | 16,091.51 |
| claude-haiku-4-5 | 1.00 | 200,000 | 5,363.84 |
| gemini-2.5-flash | 0.30 | 1,048,576 | 1,609.15 |
Topics come from an LDA model over the human-authored subset — templates removed, one document per template skeleton — fitted per language on a 2 KB prose excerpt. Topic names are not machine-generated: the model's top terms and representative documents were read by hand and labelled.
| topic | % of subset | documents | top terms | examples |
|---|---|---|---|---|
| Agentic commerce & checkout instructions | 14.14 | 12,432 | use, checkout, buyer, agents, agent, store, user, payment | |
| Local business hours & contact details | 13.39 | 11,777 | phone, contact, closed, hours, location, monday, social, media | |
| Enterprise IT, data & security | 8.59 | 7,553 | data, platform, software, management, security, solutions, cloud, company | |
| SaaS apps, features & pricing | 8.10 | 7,124 | app, free, platform, data, tools, time, features, pricing | |
| Site metadata, canonicals & sitemaps | 6.71 | 5,905 | content, site, page, use, public, llms, txt, canonical | |
| AI/API developer docs & MCP servers | 5.90 | 5,185 | api, agent, image, model, mcp, code, url, agents | |
| Hospitality, travel & bookings | 5.76 | 5,062 | hotel, travel, dental, events, booking, event, private, food | |
| Industrial equipment & trade services | 5.06 | 4,448 | company, solutions, products, industrial, repair, quality, service, equipment | |
| Design studios & creative portfolios | 4.97 | 4,373 | design, studio, brand, work, content, creative, video, build | |
| Legal, health & professional practices | 4.39 | 3,859 | research, law, firm, legal, practice, health, services, personal | |
| Digital marketing & SEO agencies | 3.68 | 3,235 | marketing, digital, services, seo, content, development, agency, web | |
| Real estate, insurance & finance | 3.64 | 3,201 | real, insurance, estate, market, property, financial, home, job | |
| Education & language training | 3.19 | 2,803 | english, india, training, languages, international, online, school, education | |
| Artist catalogues & medical care | 2.69 | 2,367 | care, artist, health, lyroverse, medical, catalog, treatment, lyric | |
| Local services (cleaning, removals) | 2.42 | 2,124 | services, service, information, page, website, business, site, canonical | |
| Manufacturing & wholesale suppliers | 2.11 | 1,856 | products, quality, game, high, games, contact, home, china | |
| Online stores, orders & shipping | 1.61 | 1,418 | product, products, store, delivery, shopify, payment, order, online | |
| Product documentation (GitBook-style) | 1.44 | 1,266 | documentation, question, gitbook, page, need, answer, query, additional | |
| Sales, CRM & messaging tools | 1.40 | 1,232 | sales, customer, community, crm, email, business, lead, support | |
| WordPress post & category dumps | 0.82 | 721 | url, published, modified, posts, categories, amp, tags, content |
| topic | % of subset | documents | top terms | examples |
|---|---|---|---|---|
| Company profiles & B2B consulting | 30.39 | 1,458 | unternehmen, deutschland, gmbh, kontakt, service, website, beratung, mail | |
| Coaching & personal brands | 13.59 | 652 | team, kunden, fragen, leben, menschen, zeit, große, unternehmen | |
| Online shops & delivery | 12.80 | 614 | eur, online, deutschland, shop, produkte, llms, website, inhalte | |
| Real-estate agents | 9.57 | 459 | immobilien, unternehmen, beratung, kunden, dienstleistungen, gmbh, bieten, service | |
| Marketing & web agencies | 8.38 | 402 | marketing, unternehmen, seo, website, agentur, content, social, design | |
| Restaurants & hotels | 5.09 | 244 | website, hotel, facebook, restaurant, uhr, kontakt, informationen, the | |
| Regional events & city portals | 4.11 | 197 | region, seiten, kinder, uhr, stadt, informationen, events, salzburg | |
| Swiss platforms & services | 3.83 | 184 | schweiz, optional, schweizer, uhr, website, innen, deutsch, plattform | |
| Retail & buying guides | 2.23 | 107 | fuer, shop, online, berlin, google, artikel, beratung, produkte | |
| Travel & outdoor in the Alps | 1.73 | 83 | hotels, outdoor, reisen, gmbh, tirol, aktivitäten, schweizer, reise | |
| Tax advisers & car dealerships | 1.60 | 77 | steuerberatung, bmw, steuerberater, unternehmen, privatpersonen, partner, standard, physiotherapie | |
| WordPress post listings | 1.35 | 65 | url, published, modified, seiten, deutsch, beiträge, tags, content | |
| Medical & aesthetic clinics (ChatGPT artefacts) | 1.33 | 64 | praxis, behandlungen, wien, behandlung, ästhetische, contentreference, oaicite, index | |
| Last-minute holiday offers | 1.08 | 52 | buchen, lastminute, angebote, hotels, urlaub, fuer, trip, ferienwohnungen | |
| Regional trades & design | 0.73 | 35 | hamburg, schleswig, holstein, online, air, uuml, baden, new | |
| Holiday flats & wellness | 0.73 | 35 | apartments, offer, personen, markdown, erhalte, line, reinigung, spirituelle | |
| Package holidays & holiday homes | 0.58 | 28 | buchen, price, urlaub, weiterlesen, trip, finde, reisen, hotels | |
| Custom manufacturing & photography | 0.44 | 21 | rahmen, bike, aluminium, main, act, münchen, stil, anfrage | |
| Privacy notices & site metadata | 0.31 | 15 | daten, beschreibung, aktuelles, kategorien, überblick, informationen, abo, website | |
| Bathroom & renovation trades | 0.13 | 6 | dusche, dresden, umbau, statt, raus, rein, bad, mainz |
| topic | % of subset | documents | top terms | examples |
|---|---|---|---|---|
| Web agencies & custom development | 19.74 | 797 | web, agence, contact, services, gestion, clients, entreprises, création | |
| Discussion-forum thread exports | 10.00 | 404 | post, patrick, url, category, created, views, utc, replies | |
| Hotels & restaurants | 8.84 | 357 | hôtel, contact, maison, restaurant, email, chambres, location, paris | |
| Content platforms & coaching | 8.67 | 350 | français, articles, contenu, contact, principales, plateforme, données, informations | |
| Building trades & call-outs | 6.69 | 270 | entreprise, services, intervention, bois, saint, zone, rcs, dépannage | |
| Cleaning companies | 5.50 | 222 | nettoyage, entreprise, réseau, econeto, saint, interventions, ligne, national | |
| Product catalogues & quotes | 5.42 | 219 | produits, livraison, ligne, entreprise, offres, url, services, service | |
| Law firms & notaries | 5.35 | 216 | droit, maître, cabinet, avocats, immobilier, téléphone, coordonnées, adresse | |
| How-to guides & blog advice | 5.05 | 204 | conseils, pratiques, guide, articles, astuces, découvrez, url, blog | |
| Insurance & Paris services | 4.19 | 169 | assurance, paris, services, saint, ligne, informations, service, contact | |
| Venue catalogues & room hire | 3.59 | 145 | catalogue, tel, optional, email, tva, personnes, rue, contact | |
| Artisan products & workshops | 3.47 | 140 | produits, paris, entreprise, catalogue, professionnels, contact, français, atelier | |
| Cosmetics & personal care | 3.12 | 126 | soins, optional, produits, language, url, paris, sections, marque | |
| SEO & technical services | 2.63 | 106 | contact, services, produits, seo, email, informations, ligne, the | |
| Heating engineers & llms.txt notes | 2.01 | 81 | contenu, contact, chauffagiste, informations, fichier, modèles, sections, llm | |
| Reviews & pillar-page structures | 1.78 | 72 | url, avis, ligne, pilier, structure, moyenne, alimentation, bonus | |
| Training courses & health advice | 1.73 | 70 | formations, cheveux, conseils, guide, formation, santé, professionnels, paris | |
| Cafes & campsites | 0.94 | 38 | café, camping, language, url, style, amour, generated, data | |
| Agency pages with ChatGPT artefacts | 0.92 | 37 | var, index, oaicite, contentreference, contact, projets, org, informations | |
| Vehicles & miscellaneous | 0.37 | 15 | utilisation, jour, fois, automobiles, pro, près, véhicules, vente |
Every grouped result carries five example documents, each with two links: the site as it is now, and #index — the record's row in the Hugging Face dataset split, which opens in the dataset viewer. The index is the durable reference; the same row is reachable in code as load_dataset("commoncrawl/llms.txt", "CC-MAIN-2026-30")["train"][index], and returns the exact bytes a figure was computed from long after the live page has changed or gone.
Source: the Hugging Face dataset commoncrawl/llms.txt, config CC-MAIN-2026-30 — WARC response records for */llms.txt and */llms-full.txt with HTTP 200 and a text/plain or text/markdown body, 584,107 documents. Analysis code: analyze.py in the llms-txt-experiments repository; every number here is produced by analyze.py aggregate from a single streaming pass over the corpus.
text/plain|markdown only. Sites that serve llms.txt as text/html, redirect, or 404 are absent — in this crawl those are the majority of all /llms.txt URL attempts./llms.txt success rate is therefore an adoption estimate for that population and not for the web at large, and the /llms-full.txt rate is not an adoption estimate at all, its sample having been enriched with hosts already known to serve the file.slot (a fifth of its matches were appointment slots, a Danish castle and a football manager), word-bounded judi (it was matching "judicial"), dropped bare xxx (it matched the Roman numeral in "XXXI Velada Musical"), and required two occurrences rather than one, which excluded a car service advertising "casino trips". Measured precision after those changes is about seven in eight.o200k_base).