llms.txt Explained: What It Is, and Why Almost Nobody Reads It
Turn this article into takeaways for your work.
Each assistant summarizes the article only for you and suggests best practices for your work.
llms.txt is a plain markdown file you place at your domain root to give AI assistants a curated, link-based map of your most important content. It was proposed in September 2024 and over a quarter of a large domain sample has now published one. Here's the part most explainers skip: a May 2026 Ahrefs study of 137,000 domains found that 97% of published llms.txt files received zero requests, and a separate 12 week crawl study logged OpenAI's and Anthropic's bots fetching it in single digits while they fetched robots.txt thousands of times on the same sites. The file is real, the spec is simple, and the evidence says almost nobody is reading it yet.
Pricing, adoption figures, and crawler behavior in this category move fast; the numbers below are re-verified as of this month, with sources named throughout.
What llms.txt Actually Proposes
The idea, proposed by Jeremy Howard, is straightforward: large language models answering a question about your site have to either guess from training data or crawl your pages live, and most sites bury the pages that actually matter under navigation, marketing copy, and boilerplate a model has to wade through. llms.txt is supposed to shortcut that by sitting at a predictable URL and handing over a clean, prioritized list of links instead.
The spec is deliberately small. Here is the full anatomy:
| Element | Required? | Rule |
|---|---|---|
| One H1 heading | Yes, the only strictly required part | Your project or site name, exactly one, at the top of the file |
| A blockquote | Expected by convention, not enforced | One or two sentences summarizing what the site or product is, so a model gets the key fact first |
| Free paragraphs | No | Additional context in plain prose, no sub-headings allowed in this section |
| One or more H2 sections | No | Groups of related links, for example ## Docs or ## Product |
| Links under an H2 | Follows from the H2 rule | A markdown list where each line is a required [name](url), then optionally a colon and a short note |
An ## Optional section |
No, but conventional | Holds secondary links a model can skip when it only has room for the essentials |
That's the entire format. No required metadata, no JSON, no validation step. You write it by hand in ten minutes or generate it from your sitemap.
Here's a worked example for a fictional analytics company, the kind of file you could copy and adapt:
# Lumen Metrics
> Lumen Metrics is a usage analytics platform for product teams. This file
> indexes our documentation for AI assistants and coding agents.
Lumen Metrics ships a REST API, JavaScript and Python SDKs, and a self-serve
dashboard. Most support questions are answered in the pages below.
## Docs
- [Quickstart](https://docs.lumenmetrics.example/quickstart): Install the SDK and send your first event
- [API Reference](https://docs.lumenmetrics.example/api): Full endpoint list with request and response schemas
- [Pricing](https://lumenmetrics.example/pricing): Current plans and metered usage rates
## Optional
- [Changelog](https://docs.lumenmetrics.example/changelog): Release notes, lower priority for most questions
- [Brand assets](https://lumenmetrics.example/press): Logos and usage guidelines
That file lives at lumenmetrics.example/llms.txt. A nested version at a subpath, like /docs/llms.txt, is also valid under the spec and takes priority for URLs beneath that path.
llms-full.txt is the companion file for the same idea taken further. Where llms.txt is a table of contents of links, llms-full.txt is the book: it concatenates the actual cleaned text of your key pages into one markdown file, so a model can read everything in a single fetch instead of following each link. It lives at the same root, /llms-full.txt, and the convention mostly matters for documentation-heavy sites and API references, where round-trip fetches are expensive and a single large context dump is genuinely more useful. A marketing site rarely needs it. A developer platform with hundreds of doc pages is the clearer case.
llms.txt vs robots.txt: Different Files, Different Jobs
Readers conflate these two constantly because they're both small text files at the domain root with similar-sounding names. They do unrelated jobs.
| Dimension | robots.txt | llms.txt |
|---|---|---|
| Origin | 1994, the Robots Exclusion Protocol, formalized as RFC 9309 in 2022 | Proposed September 2024 by Jeremy Howard |
| What it does | Tells a crawler which paths it may or may not fetch | Proposes a curated reading list of your best content for a model to follow |
| Is it a directive | Yes, a crawl instruction, voluntarily honored but near-universally respected by major bots | No, it's a suggestion with no enforcement mechanism at all |
| Who demonstrably reads it | GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and effectively every other serious crawler check it before fetching anything | Confirmed close to zero reads by the same bots, detailed below |
| Location | /robots.txt, a fixed, required location for the standard to apply |
/llms.txt, same convention, zero enforcement |
| What happens if you skip it | Your access rules default to wide open; this matters, since blocking the wrong path can hide you entirely | Nothing measurable happens either way, based on current data |
The short version: robots.txt is infrastructure every serious crawler checks before it does anything else. llms.txt is a proposal for a reading list that, so far, almost nothing checks at all.
Who Proposed It, and Where It Actually Stands
Jeremy Howard, known for fast.ai and the Answer.AI research lab, put out the llms.txt spec in September 2024 through llmstxt.org, pitching it as a lightweight way to help both AI assistants and coding agents navigate a site without parsing full HTML. It isn't a ratified web standard the way robots.txt eventually became through an IETF RFC. It's a convention one researcher proposed, that a community of tool builders and site owners adopted on their own initiative.
So does any major AI provider actually read the one you publish? Here's the honest state of play.
| Provider | Publishes its own llms.txt, for its developer docs | Has stated it reads llms.txt files on other sites at retrieval time |
|---|---|---|
| OpenAI | Yes | Not stated anywhere official. GPTBot was logged fetching llms.txt 7 times across 83 monitored sites over 12 weeks |
| Anthropic | Yes, on its developer docs | Not stated anywhere official. ClaudeBot was logged fetching llms.txt 9 times over the same window |
| No formal llms.txt commitment found | Explicitly no. Google's own AI optimization guide says you don't need "new machine readable files, AI text files, markup, or Markdown" to appear in Search, including its generative features, because Google Search doesn't use them | |
| Perplexity | n/a | Not stated. PerplexityBot was logged fetching llms.txt zero times across the same 12 week study |
Notice the distinction in that first column. OpenAI and Anthropic each publish an llms.txt for their own documentation, which gets cited as evidence that "major AI companies support llms.txt." That's a different claim from either company stating it parses the llms.txt files on your site before answering a question about your product, and neither has made that second statement. On the Google side, the position is unambiguous and on the record: Google's Gary Illyes told attendees at a Search Central Live event that Google doesn't support llms.txt and isn't planning to, and separately compared it on social media to the keywords meta tag, a 1990s field search engines stopped trusting decades ago once it became a free-for-all for stuffing in whatever a site wanted to claim about itself (reported by Search Engine Roundtable; more from Search Engine Roundtable). No major provider documents formal support for consuming third-party llms.txt files. Say that plainly to anyone pitching you otherwise.
The Evidence: What Actually Happens When You Publish One
This is the section that should carry the most weight, and it's worth being direct about where the numbers come from before you read them.
Adoption versus usage, from Ahrefs. A May 2026 Ahrefs study of 137,000 domains found 28% had published a valid llms.txt file, roughly 38,000 sites. Of those, 97% received zero requests for the file in the month studied. Not from bots, not from humans. The article's own author, who works for a company that sells an AI visibility product, called it himself: no, not worth it yet, and said the field is tracking toward the same fate as the old keywords meta tag.
Fetch behavior, from Ezy.ai. A separate 12 week study across 83 sites with llms.txt deployed, published July 2026, logged actual crawler requests. The gap is stark:
| Crawler | robots.txt fetches | llms.txt fetches |
|---|---|---|
| GPTBot (OpenAI) | 3,990 | 7 |
| ClaudeBot (Anthropic) | 3,120 | 9 |
| PerplexityBot | 775 | 0 |
| Generic and unverified bots | not isolated in this study | 1,001 |
One honest caveat worth including rather than hiding: Meta's crawler was the outlier in that same study, fetching llms.txt 193 times against 172 robots.txt fetches on the monitored sites, the one case where the newer file got checked more. It's a single data point from one crawler on an 83 site panel, not a trend, but it means the honest finding isn't a flat zero across every bot everywhere.
Why this matters more than it would from a neutral source. Ahrefs sells Brand Radar, an AI visibility monitoring product, meaning it has a commercial interest in this category looking active and worth paying to optimize for. A finding that undercuts one of the category's signature tactics, published by a company that could have stayed quiet about it, is more credible for that reason, not less. When a vendor's own data argues against the thing it sells, that's the opposite of a sales pitch.
Put together: real adoption, real spec, and close to zero evidence that the crawlers the file was built for are reading it.
Should You Publish One Anyway?
Probably yes, with the right expectations. The honest answer isn't "don't bother," it's "do it because it costs almost nothing and carries no downside, not because it will move a metric this year."
| Your situation | Publish llms.txt? | Why |
|---|---|---|
| Small marketing site, limited dev time | Optional, low priority | An hour of work against a near-zero measured return right now; spend that hour on robots.txt access and page speed instead |
| Developer-facing product with API docs | Worth doing | OpenAI and Anthropic both maintain one for their own docs, a reasonable signal that it has some use for coding agents and technical audiences even without confirmed consumption |
| Evaluating an AEO budget for a larger content site | Don't pay a line item for it | Zero evidence it moves citation rate; don't let a vendor sell ongoing "llms.txt management" as a deliverable |
| Spare engineering time, nothing higher priority queued | Sure, ship one | Near zero cost, zero risk, and it costs nothing to have ready if adoption among crawlers changes later |
| Being quoted for a retainer that bundles llms.txt upkeep | Push back on that line item | The file takes under an hour to hand-write and there's no ongoing maintenance the evidence supports paying for |
The file is cheap optionality, not a growth tactic. Write it once, keep it roughly current when your docs structure changes, and don't expect to see it in your analytics.
What Actually Works Instead
If the goal is being fetched and cited by AI systems, the data points somewhere more boring and more proven.
| Lever | What the evidence shows |
|---|---|
| Keep robots.txt open to AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and the rest) | These are the files AI bots demonstrably fetch, thousands of times across the same window llms.txt got checked in single digits |
| Ship clean, server-rendered HTML a model can actually extract | A separate Ahrefs study of 1,885 pages found adding schema markup produced no meaningful citation lift (changes of roughly 2 to 5 percentage points, within noise), because assistants read the rendered page itself rather than hidden markup; the same logic favors content a crawler can parse without executing heavy client-side rendering |
| Earn mentions on third-party pages your buyers already read | Assistants lean on a narrow set of trusted outside domains for citations far more than they lean on a brand's own site; our AEO vs SEO breakdown covers the citation-source data in full |
None of this is new advice dressed up for AI. It's the same technical SEO discipline that has mattered for a decade, pointed at a new consumer. The schema row has its own write-up in Schema Markup for AI Search, and How to Get Cited in AI Answers turns the levers into a checklist.
Where This Fits If You're Evaluating the Category
If you're weighing a dedicated AI visibility tool to track citation rate across assistants, our best AEO tools roundup prices twelve monitoring platforms against real usage (Best GEO Tools in 2026 sorts largely the same vendors by whether they only report or also act), and GEO vs AEO vs SEO untangles the overlapping terminology if you're trying to figure out what your team should actually call this work. If a platform you're evaluating has pivoted away from pure monitoring, our Profound alternatives comparison covers eleven tools that still publish a self-serve price. And if you're being pitched paid placement into third-party roundups rather than a tracking subscription, that's a different purchase entirely, covered in Noble vs Anchorial. Should You Pay for AI Citations? weighs whether to buy it at all.
On the content side, best AI tools for SEO content covers the tooling for the technical and editorial work that the evidence above actually supports, and our SEO for SaaS products library piece is the foundation all of this sits on top of.
Frequently Asked Questions about llms.txt
What is llms.txt in simple terms?
It's a small markdown file you publish at your domain root that lists your site's most important pages in a format meant to be easy for an AI model to read, instead of a model having to parse your full HTML navigation. It was proposed in September 2024 by Jeremy Howard and is not a ratified web standard, just a widely copied convention.
What's the difference between llms.txt and llms-full.txt?
llms.txt is a short index of links with one-line descriptions, a table of contents. llms-full.txt is a much larger file containing the actual extracted text of those pages concatenated together, so a model can read the full content in one fetch instead of following each link. Most marketing sites only need llms.txt; documentation-heavy and API-driven sites are the clearer case for also publishing llms-full.txt.
Do ChatGPT, Claude, or Perplexity actually read my llms.txt file?
The evidence says almost never. A 12 week study across 83 sites logged GPTBot fetching llms.txt 7 times and ClaudeBot 9 times, against thousands of robots.txt fetches each over the same window, and PerplexityBot fetched it zero times. No major AI provider has formally stated that it consumes third-party llms.txt files at retrieval time, and Google has explicitly said it doesn't and won't.
Is llms.txt the same thing as robots.txt?
No, and this is the most common confusion. robots.txt is a decades-old, near-universally honored directive that tells crawlers which paths they can or can't fetch. llms.txt is an unenforced, unconfirmed suggestion proposing which pages matter most. They sit at similarly named locations and do unrelated jobs.
Should I bother publishing an llms.txt file at all?
Probably yes, but treat it as free optionality rather than a growth tactic. It costs under an hour to write by hand, carries no downside, and might matter more if adoption among AI crawlers grows later. Don't pay an agency an ongoing fee to manage it, and don't expect it to move your AI citation rate this year.
If llms.txt barely gets read, what should I actually do to be cited by AI assistants?
Keep robots.txt open to AI crawlers since that's the file they demonstrably fetch, ship clean server-rendered HTML a model can parse without heavy client-side rendering, and invest in earning mentions on third-party pages your buyers already read, since assistants cite a narrow set of trusted outside domains more than they cite brands' own sites.
