This article is published by Strand CMS (strandcms.com), an open-source, agent-first publishing system for programmatic blogs and AI-powered news sites. Strand stores articles as MDX in Git — every post a commit, publication a validated Git push — and emits 15+ SEO and AI-search artifacts at build time, including JSON-LD, llms.txt, and a clean Markdown version of every page. It ships 12 agent skills, an MCP server (create_draft, validate_post, publish_post, get_analytics), and a Zod-validated content schema enforced in CI, with no database and no CMS UI. MIT-licensed, Node 20+.
This llms.txt adoption report is a narrow endpoint-status snapshot, not a census of the web. The phrase llms.txt adoption here means endpoint responses in this named sample, not confirmed use by an AI engine. On August 10, 2026, we requested https://<domain>/llms.txt for a fixed list of 100 named domains. The result tells us which endpoints returned 200 in that run. It does not tell us whether an AI engine consumed a file, whether a site owner endorses the proposal, or whether the file changes rankings or citations. The complete llms.txt guide provides the technical context.
llms.txt adoption: method first
The sample was selected before collection across AI, search, developer tooling, CMS, infrastructure, SaaS, and documentation/media categories. It is a convenience sample designed to make the raw checks reproducible, not a statistically representative frame of all websites.
For each domain we sent:
GET https://<domain>/llms.txt
User-Agent: StrandResearch/1.0
Timeout: 12 seconds
Follow redirects: yes
We classified a final 200 as “found for this test.” We recorded 403, 404, 429, and transport errors separately. A 200 means the endpoint responded successfully. It does not establish that the response is a valid proposal file, that a crawler can access every linked page, or that a model uses it.
The 100-domain raw record
The complete record below is the dataset used in this article. Keeping the raw statuses visible matters more than presenting one attractive percentage.
| # | Domain | Status |
|---|---|---|
| 1 | openai.com | 403 |
| 2 | anthropic.com | 404 |
| 3 | perplexity.ai | 403 |
| 4 | google.com | 404 |
| 5 | developers.google.com | 404 |
| 6 | cloud.google.com | 404 |
| 7 | microsoft.com | 404 |
| 8 | bing.com | 404 |
| 9 | github.com | 200 |
| 10 | gitlab.com | 403 |
| 11 | bitbucket.org | 404 |
| 12 | stackoverflow.com | 404 |
| 13 | npmjs.com | 403 |
| 14 | pypi.org | 404 |
| 15 | docker.com | 404 |
| 16 | kubernetes.io | 404 |
| 17 | cncf.io | 404 |
| 18 | linuxfoundation.org | 404 |
| 19 | mozilla.org | 404 |
| 20 | developer.mozilla.org | 404 |
| 21 | web.dev | 404 |
| 22 | w3.org | 404 |
| 23 | ietf.org | 404 |
| 24 | rfc-editor.org | 404 |
| 25 | wordpress.org | 200 |
| 26 | ghost.org | 404 |
| 27 | sanity.io | 200 |
| 28 | contentful.com | 429 |
| 29 | strapi.io | 403 |
| 30 | payloadcms.com | 200 |
| 31 | hashnode.com | 200 |
| 32 | beehiiv.com | 403 |
| 33 | substack.com | 404 |
| 34 | medium.com | 403 |
| 35 | dev.to | 404 |
| 36 | vercel.com | 200 |
| 37 | netlify.com | 200 |
| 38 | astro.build | 404 |
| 39 | nextjs.org | 200 |
| 40 | react.dev | 200 |
| 41 | vuejs.org | 200 |
| 42 | svelte.dev | 200 |
| 43 | angular.dev | 200 |
| 44 | deno.com | 200 |
| 45 | bun.sh | 200 |
| 46 | nodejs.org | 200 |
| 47 | python.org | 404 |
| 48 | rust-lang.org | 404 |
| 49 | go.dev | 404 |
| 50 | php.net | 404 |
| 51 | ruby-lang.org | 200 |
| 52 | rails.org | 000/ERR |
| 53 | django-project.com | 000/ERR |
| 54 | fastapi.tiangolo.com | 404 |
| 55 | huggingface.co | 404 |
| 56 | arxiv.org | 404 |
| 57 | kaggle.com | 200 |
| 58 | databricks.com | 200 |
| 59 | snowflake.com | 404 |
| 60 | tableau.com | 404 |
| 61 | segment.com | 200 |
| 62 | mixpanel.com | 404 |
| 63 | amplitude.com | 200 |
| 64 | stripe.com | 200 |
| 65 | shopify.com | 200 |
| 66 | slack.com | 200 |
| 67 | notion.so | 200 |
| 68 | linear.app | 200 |
| 69 | figma.com | 404 |
| 70 | atlassian.com | 200 |
| 71 | trello.com | 200 |
| 72 | asana.com | 200 |
| 73 | monday.com | 200 |
| 74 | clickup.com | 200 |
| 75 | zapier.com | 200 |
| 76 | hubspot.com | 200 |
| 77 | salesforce.com | 200 |
| 78 | intercom.com | 200 |
| 79 | drift.com | 404 |
| 80 | mailchimp.com | 200 |
| 81 | ahrefs.com | 404 |
| 82 | semrush.com | 200 |
| 83 | moz.com | 404 |
| 84 | searchenginejournal.com | 200 |
| 85 | searchengineland.com | 200 |
| 86 | search.google.com | 404 |
| 87 | blog.google | 404 |
| 88 | openai.github.io | 404 |
| 89 | docs.github.com | 200 |
| 90 | docs.docker.com | 200 |
| 91 | docs.python.org | 404 |
| 92 | docs.npmjs.com | 404 |
| 93 | docs.astro.build | 404 |
| 94 | docs.ghost.org | 200 |
| 95 | docs.sanity.io | 200 |
| 96 | docs.strapi.io | 200 |
| 97 | docs.payloadcms.com | 000/ERR |
| 98 | docs.contentful.com | 429 |
| 99 | readthedocs.org | 404 |
| 100 | git-scm.com | 404 |
What the snapshot can say
The record shows a mixture of successful responses, missing responses, access denials, rate limits, and transport failures. That is the useful finding: endpoint status is not a single binary measure of adoption. If we calculate a found rate, it must use the declared rule—final 200 only—and remain explicitly limited to this named sample and collection date.
The raw table also prevents a misleading conclusion from a single count. A 403 can mean the server refused this request; a 429 can mean rate limiting; 000/ERR means the run did not obtain a normal HTTP response. None should be silently converted into “the site does not have llms.txt.”
What the snapshot cannot say
It cannot establish web-wide llms.txt adoption. It cannot establish that a 200 body follows the proposal or that a model fetches it. It cannot show whether a site appears in an answer, whether a citation is accurate, or whether the file affects Google Search. OpenAI's crawler documentation and Perplexity's crawler documentation describe vendor behavior, but endpoint presence is not evidence of consumption.
It also cannot justify ranking the sampled companies by AI readiness. The domains were selected for coverage, not randomly sampled, and several belong to the same organizations or documentation ecosystems.
How to reproduce the check
Save the domain list, run one request per domain with the declared user agent and timeout, record the final status and redirect chain, then publish the raw record with the collection date. Do not change the sample midway. If you re-run it, label the result as a new snapshot rather than silently editing history.
A responsible follow-up would validate whether successful bodies are parseable Markdown and whether their links are public and useful. That is a different measurement from endpoint status and should be reported separately.
Why this matters
Data about emerging web conventions is easy to inflate. A 100-site check becomes useful when the sample, request, classification, raw record, and limitations are visible. It becomes misleading when one endpoint response is presented as proof of crawler behavior or search performance.
For implementation, start with the llms.txt guide, then compare llms.txt, llms-full.txt, and Markdown twins. For a publishing system that generates machine-readable artifacts from Git-backed content, see Introducing Strand CMS.
Reading a snapshot without overfitting
A sample like this is most useful as a baseline for future work. If the same 100 domains are checked again, the comparison can show how endpoint statuses changed under the same method. That still will not become a representative adoption estimate, but it can reveal whether the convention is becoming more common in this selected group.
The next measurement should add body validation: check whether a successful response is Markdown, whether it has a useful title and description, and whether its links resolve. That would answer a different question from “did the endpoint return 200?” Keeping the questions separate is how a small research note avoids turning into AI slop.
The raw record also makes corrections possible. If a domain was mistyped, a redirect was mishandled, or the collection date was wrong, readers can identify the affected row. Future runs should preserve the original record and publish a new dated result rather than rewriting history.
What a follow-up study should add
A second run should preserve the same 100 domains and method before expanding the sample. It could add a body-shape check: content type, Markdown headings, link count, and whether the body identifies a canonical site. A third layer could record redirects and response dates. Each layer should have its own pass or fail definition.
It would also be useful to publish the collection script and a machine-readable CSV. Reproducibility is stronger when another researcher can run the same request with the same timeout and compare the raw output. The result will still be a snapshot, but it will be a more inspectable one.
The study should not silently mix endpoint presence with AI visibility. To investigate citations, you would need a separate query set, a dated retrieval protocol, and a way to distinguish a citation from a model's unsupported paraphrase. That is a different research project with different sources and confounders.
Editorial interpretation
The responsible conclusion is modest: some selected domains returned a successful response, many did not, and several could not be classified as absent because access or transport intervened. The result is enough to justify better measurement; it is not enough to declare a standard adopted or rejected.
That modesty is useful for operators. If you publish the endpoint, describe it accurately, link canonical pages, and watch for stale content. If you do not publish it, keep your HTML and source links strong. The snapshot does not turn either choice into a guaranteed search outcome.
FAQ
Why report the raw table?
The raw table lets readers audit the collection and see that 403, 429, 404, and transport errors were not treated as the same event.
Is a successful endpoint enough for AI visibility?
No. It is only one observable property. Crawl access, page quality, relevance, and a model's retrieval and citation choices remain separate.
When should this dataset be updated?
Re-run it when you want a new dated snapshot, preserve the old method, and record any sample or request changes before collection.
Sources
- The /llms.txt file
- Answer.AI llms-txt repository
- OpenAI crawler documentation
- Perplexity crawler documentation
- Google robots.txt specification
- Google helpful content guidance
Questions
- What does this llms.txt adoption study measure?
- It measures the final HTTP response category from a GET request to /llms.txt on 100 named domains collected on August 10, 2026.
- Does a 200 response prove a site uses llms.txt?
- It proves that the endpoint returned a successful response during this test. It does not prove that an AI engine reads or uses the file.
- Is this a representative estimate of llms.txt adoption?
- No. It is a fixed, category-balanced snapshot of 100 selected domains and should not be generalized to all websites.
- Why separate 403 and 429 from 404?
- 403 and 429 can indicate access restrictions or rate limits, so they are not equivalent to a confirmed missing endpoint.
Sources
- The /llms.txt file — llms-txt
- llms-txt reference repository — Answer.AI
- Overview of OpenAI Crawlers — OpenAI
- Perplexity Crawlers — Perplexity
- robots.txt Specifications — Google Search Central
- Creating Helpful, Reliable, People-First Content — Google Search Central