llms.txt Explained: The Complete Guide
llms.txt is a Markdown file at the root of a website (/llms.txt) that gives large language models and AI agents a short summary of the site and a curated list of links to its most useful pages, so they can understand it without wading through navigation, scripts and ads. This guide covers the exact llms.txt format with a real example, how it differs from robots.txt and llms-full.txt, how to create one, and what Google and AI assistants actually do with it in 2026.
What is llms.txt?
Web pages are built for people. A model reading one has to dig the actual content out from menus, cookie banners, related-post widgets and inline JavaScript, often inside a tight context window where every wasted token pushes useful text out. llms.txt is a small fix for that: one plain Markdown file that says what the site is and points straight at the pages worth reading, each with a one-line note on what it covers.
It was proposed by Jeremy Howard at llmstxt.org in September 2024, and the proposal was revised as version 2 in August 2026. It isn't an IETF or W3C standard, but it has spread quickly: documentation platforms now generate one automatically, and OpenAI, Anthropic and Google (for the Gemini API docs) each publish one. The name is a deliberate nod to robots.txt, but the two files do very different jobs, covered below.
The llms.txt file format
An llms.txt file is ordinary Markdown with a fixed order. Only the first part is required:
- An H1 with the site or project name (
# Acme). The single required element. - A blockquote summary (a line starting with
>) holding the key facts a model needs to make sense of everything else. - Optional details: paragraphs or bullet lists with extra context, but no headings.
- H2 sections of links, each item written as
- [Page name](https://url): short note. The note after the colon is optional but is what lets a model choose a page without opening it. - An "Optional" section for secondary links. In version 1 this name told tools they could drop those links to save space; version 2 removed that special behaviour, so it's now simply a clear place for lower-priority pages.
Version 2 also settled where the file can live. The main one sits at the root, /llms.txt, but a subfolder can have its own: /docs/llms.txt covers the pages under /docs/, and when two files apply, agents should use the more specific one.
A real llms.txt example
This is the start of Frost Rank's own file at frostrank.com/llms.txt, copied from the live site (the "…" marks lines cut for length):
# Frost Rank
> Frost Rank is a free suite of fast, no-signup SEO and developer tools — schema generators, meta tag builders, redirect checkers, formatters and more — plus in-depth guides on the technical SEO and developer concepts behind them.
## Tools
- [JSON-LD Schema Generator](https://frostrank.com/tools/json-ld-schema-generator): Generate valid structured data for Person, Organization, Product, FAQ and more — no hand-written JSON-LD.
- [Meta / OG Tag Generator](https://frostrank.com/tools/meta-og-tag-generator): Build accurate title, description, Open Graph and Twitter Card tags with live previews.
…
## Guides
- [JSON-LD Structured Data: The Complete Guide](https://frostrank.com/guides/json-ld-structured-data-guide): What JSON-LD structured data is, which schema types are worth your time, how Google really uses it, and the mistakes that get pages penalized or ignored.
…
Three choices in it are worth copying. The summary names what the site is and what it isn't (free, no signup) in one sentence. Every link is an absolute https:// URL, because a model often reads the file on its own with no page to resolve /tools/… against. And every link carries a description, so a model asked about schema markup can go straight to the right page. The file is generated from the same configuration that builds the site's menus and sitemap, so a new tool appears in it the moment it launches.
Markdown versions of your pages
llms.txt is the table of contents; the proposal also suggests offering each page itself as clean Markdown. Version 2 accepts two URL forms, the page URL with .md appended (page.html.md) or with the extension replaced (page.md), and adds standard link relations so an agent can find both without guessing: rel="alternate" type="text/markdown" points to a page's Markdown version, and rel="describedby" points to the llms.txt file that covers it. As an HTTP header, that looks like:
Link: </docs/page.html.md>; rel="alternate"; type="text/markdown", </docs/llms.txt>; rel="describedby"
Separate .md files aren't the only way to do this. Frost Rank uses content negotiation instead: when a request's Accept header asks for text/markdown, the same URL returns the page converted to Markdown, with the navigation, footer, scripts and structured data stripped out first. One URL serves both audiences, and there are no duplicate files to keep in sync.
llms.txt vs robots.txt vs sitemap.xml
All three are plain files at the root of a site, and all three are read by machines, which is why they get mixed up. They answer different questions:
| robots.txt | sitemap.xml | llms.txt | |
|---|---|---|---|
| Question it answers | What may crawlers fetch? | Which URLs exist? | What should an AI read first? |
| Audience | All crawlers | Search engines | LLMs and AI agents |
| Format | Plain-text directives | XML | Markdown |
| Scope | Every path, by rule | Every indexable URL | A curated shortlist |
| Can block access? | Yes | No | No |
| Status | Standard (RFC 9309) | Standard (sitemaps.org) | Community proposal |
The practical consequence: llms.txt can't keep an AI crawler out, and listing a page in it doesn't let a crawler in. Access for bots like GPTBot, ClaudeBot or Google-Extended still lives in robots.txt, and the robots.txt and XML sitemaps guide covers how those rules are matched.
What about llms-full.txt?
llms-full.txt is a companion convention, popularised by documentation platforms rather than the llms.txt proposal itself. Instead of links, it contains the full text of every page in one Markdown file, so a tool can load an entire documentation set in a single request. It works well for a compact product's docs that fit comfortably in a model's context window. For a large site it backfires: the file can run to millions of characters, far past what a model can hold at once, and the tool reading it has to cut it down anyway. Publish llms.txt first, and add llms-full.txt only if your docs are small enough to make it useful.
Does llms.txt help SEO?
Not for Google rankings, and Google has said so plainly. Its documentation on AI features and your website states that you don't need to create new machine-readable files, AI text files or markup to appear in AI Overviews or AI Mode, and its May 2026 post on optimizing for generative AI in Google Search lists llms.txt among the tactics you can skip. Those features pull from the regular Search index, so the same things that earn rankings, like crawlable pages, helpful content and clean technical SEO, decide whether you're cited.
The same month, Chrome added an llms.txt check to Lighthouse's new Agentic Browsing category, which looks for the file at the domain root. The two messages don't contradict each other: Search says the file won't change rankings, while Lighthouse treats it as something that helps AI agents navigate a site on a user's behalf.
That's the honest summary of where llms.txt helps today:
- Worth doing for API and developer documentation, SaaS products, and any site people point AI coding assistants or agents at. A clean index spares those tools from guessing which pages matter.
- Cheap insurance for content sites, blogs and online stores. It takes minutes and does no harm, but don't expect traffic from it.
- Not a ranking lever anywhere. If someone sells llms.txt as a way to rank in AI Overviews or ChatGPT, they're overselling it.
How to create an llms.txt file
-
1
Pick the pages that matter
List the 10 to 50 pages that best explain your product or site: docs, pricing, key guides, policies. Leave out tag archives, login pages and anything thin.
-
2
Write the summary
Start the file with an H1 holding your site name, then a one- or two-sentence blockquote that states what the site is and who it is for.
-
3
Group links under H2 sections
Add sections such as Docs, Guides or Products, and list each page as "- [Name](https://full-url): what it covers". Put secondary pages in a final section named Optional.
-
4
Upload it to the root of your site
Save the file as llms.txt and serve it at https://yoursite.com/llms.txt as plain text, with a 200 status and no redirect.
-
5
Validate the live file
Fetch the live URL with a validator to confirm it returns the file itself rather than an HTML 404 page, and that every link is absolute and working.
The llms.txt Generator pre-fills steps 1 to 3 from your sitemap, then validates the live file in step 5: it checks the status code and Content-Type, flags a missing H1 or summary, and catches relative or duplicate links.
Common llms.txt mistakes
- The server returns a web page instead of the file. Many sites answer unknown paths with a styled 404 or, on single-page apps, the homepage, both with a 200 status. An AI tool then reads your HTML shell, not your summary. Check that
/llms.txtreally returns plain text; the HTTP Header Checker shows the status and Content-Type. - Dumping the whole sitemap. Hundreds of undescribed links recreate exactly the noise llms.txt is meant to remove. Curate.
- Relative links.
/docs/apimeans nothing to a model that loaded the file directly. Use full URLs. - Links to blocked or redirected pages. If robots.txt disallows the crawler, or a link bounces through redirects, the page never gets read. List final URLs that the bots you care about may fetch.
- Putting private information in it. The file is public. Don't list staging URLs, internal tools or anything you wouldn't put on your homepage.
- Letting it go stale. A file written once by hand drifts out of date as pages move. Generate it from your site's own data where you can.
Frequently asked questions
No. It's a community proposal by Jeremy Howard, first published in September 2024 and revised as version 2 in August 2026, not an IETF or W3C standard. See where llms.txt came from.
No. Google says you don't need AI text files like llms.txt to appear in AI Overviews or AI Mode, which use the regular Search index. See what Google and Chrome have said.
They can read it whenever someone, or an agent acting for them, fetches it, for example when you ask an assistant to read yoursite.com/llms.txt or a coding agent looks up your API docs. That is where the file earns its keep. None of the major assistants documents llms.txt as a signal for deciding which websites to cite in everyday answers, so don't expect it to change how often you are mentioned.
No. robots.txt controls which crawlers may fetch what; llms.txt controls nothing and is only a reading list, so AI crawler access still belongs in robots.txt. See the three files compared.
Yes. A file in a subfolder, like /docs/llms.txt, covers the pages under it, and agents use the most specific file that applies. See the file format and where files can live.
llms.txt is an index of links; llms-full.txt is a separate convention that puts the full text of every page into one file, which only suits small documentation sets. See when llms-full.txt makes sense.
Whenever the pages it links to change: a new product, a renamed docs section, a removed page. Stale links are worse than none, because an agent that follows a dead link wastes its context on a 404. Generating the file from the same data that builds your site, as Frost Rank does, keeps it current without manual edits.
Based on running llms.txt and Markdown-for-agents in production on this site, and checked against the llms.txt v2 proposal and Google's and Chrome's own published guidance rather than second-hand summaries.
Was this guide helpful?
Rate it and leave a comment — takes 10 seconds, and genuinely helps decide what to fix or build next.