Skip to content
Search & AI

llms.txt: the file that introduces your site to AI

llms.txt tells AI what to read on your site. The work that matters is deciding which pages represent your brand

Type your site’s address, add /llms.txt, and hit enter: your server might return a file you didn’t even know existed.

At the top you’ll find the domain name and a couple of introductory lines, below that a long list of links grouped by section. If nobody at your company remembers writing it, a plugin or your CMS generated it when some option got switched on, and it’s been presenting your brand with wording you never approved ever since.

llms.txt is known for telling AI systems which content to cite when someone asks about you. Chances are none of them has actually read it yet.

What llms.txt is

llms.txt is the file you use to tell programs built on language models which content on your site deserves their attention, and in what order to read it. It’s a Markdown document, the syntax that marks headings and lists with a handful of symbols while staying readable to both a person and a piece of software. At the top goes the project name, the only part the spec requires; below that you can add a short summary, useful for interpreting the rest, and the resource links grouped by topic. Each entry carries the resource’s title, its address, and optionally a note on what’s behind the link.

The programs you’re addressing are AI agents, software that opens pages and documentation to complete a task assigned by a person. An agent is the coding assistant a developer asks to connect your service to the software they’re building; it’s also the chatbot that checks a subscription’s terms before answering a question. They arrive with a task already defined, and they read your site based on that.

The scope of the declaration is set by the folder where you publish the file. At the domain root it covers the whole site; inside a path it applies only to the URLs beneath it, so an llms.txt in /docs/ describes the documentation while the rest of the catalog stays outside its scope. Where multiple files overlap, an agent follows the more specific one. You can therefore describe just one section, even when you don’t control the domain root.

llms.txt isn’t a W3C standard or an IETF RFC, and no program is required to implement it. Jeremy Howard published the spec on September 3, 2024, on the llmstxt.org domain, on behalf of Answer.AI and fast.ai, the research lab and open source project he works on language models with, and the document presents itself as an open proposal. The spec doesn’t define how content gets selected, so two plugins can generate different files while still following the same format.

On August 10, 2026, Howard published version 2, which formalizes the per-path scope, introduces links for finding the Markdown copies, and reduces “Optional” to a convention for secondary links. Any guide that hasn’t been updated since that date needs checking against the current version.

What it costs an AI agent to read your site

An agent that lands on your site has no reason to crawl it indiscriminately. The model behind it works within a context window, the amount of text it can keep in mind at once, and your pages’ HTML eats into that with menus, scripts, banners, and cookie notices before it even reaches the content. Extracting clean text from that code is still an imprecise process, and the tokens spent on throwaway code are compute time you don’t get back.

Prova subito AI Visibility di SEOZoom

An llms.txt offers a leaner solution: a document light enough to fit entirely in the context window, that lists resources instead of containing them and leaves the details behind links, fetched only when needed. The spec leaves the choice of addresses and their order up to whoever publishes the file.

What sets it apart from robots.txt and the XML sitemap

robots.txt shares its location and file extension with llms.txt, but governs a different matter. In robots.txt you declare which paths a crawler is allowed to request, and programs that follow the protocol read those rules before asking your server for anything. An llms.txt comes into play later, once an agent is already looking for material to complete a task. A section under Disallow therefore stays off-limits to anything that respects the directive, even if it shows up among the links in the summary. ChatGPT-User and Perplexity-User are a declared exception: their operators state that when a fetch originates from a direct user request, those rules may not apply.

The «AI sitemap» nickname has been around since the start. But you hand XML sitemaps to search engines so they can discover the URLs you want crawled, and they list everything. An llms.txt only points to the documents that matter for a given scope, and you can even include external resources when they help explain your own. Howard adds a practical reason: a sitemap typically covers more documents than any context window could hold.

A URL listed in llms.txt keeps the canonical tag, robots directives, and indexing conditions it already had. It gets into Google if and when it would have gotten in anyway, and it stays off-limits to any crawler that respects a Disallow.

How much llms.txt actually matters, and who really opens it

llms.txt plays a much narrower role than its reputation suggests. It does give agents a compact map of resources you want to be easy to find, with notes on what they contain and what they’re for, but it doesn’t control crawling, which is still the job of robots.txt, and it doesn’t assign priority to pages in search systems. It organizes a selection and makes it readable to a program that decides to consult it.

So the value lies mostly in the clarity it adds. For documentation, APIs, and large resource collections, it can cut down the work needed to find the right document. For blogs, ecommerce sites, and corporate websites, the benefit is less clear. On top of that, we don’t currently have solid evidence that having an llms.txt increases your odds of being picked or cited in generative answers. Publishing it adds a possible path to your content, not a ranking signal or a visibility guarantee.

That doesn’t make it useless. If your sitemap, page accessibility, structured data, and robots.txt are already in order, putting together llms.txt forces you to spell out which resources actually serve a purpose for an agent and how you want to describe them. If the file already exists because your CMS generated it, checking it is even simpler: it’s worth confirming that the selection and the descriptions still match your site. If you have to build it from scratch, priority depends on the use you can reasonably expect and, above all, on what you’re actually able to observe.

On that last point, we do have some data. In May 2026 Ahrefs analyzed requests to llms.txt across 137,210 domains connected to its own analytics service (sites more technical and more SEO-aware than the average on the web, as the authors themselves pointed out): 38,360 of them published a valid file, and 97% didn’t get a single request during the month. A logged request, moreover, documents a fetch and doesn’t prove the file was actually read or used, so these figures describe an upper bound on real use, not actual use. On domains that don’t publish one, no AI system went looking for it.

96% of requests come from bots that have little to do with AI at all: SEO tools, generic crawlers, validators, services that study the adoption of the format itself. Identified AI systems account for 19.5% of requests, and among the main categories, agents make up 10.5%, training crawlers 5.3%, and the bots that fetch sources for generative answers – OAI-SearchBot and PerplexityBot – just 1.1%.

Meanwhile, even in its own guidance on optimizing for generative AI features, Google states that showing up doesn’t require AI-specific files, extra markup, or Markdown versions, because Search doesn’t use them, a position it has repeated several times. Keeping the file for other services is still fine, and it neither helps nor hurts visibility. AI Overviews and AI Mode work off the index Googlebot builds by visiting your pages, and a page can be considered for those answers once it’s indexed and eligible for a snippet.

Lighthouse cerca invece /llms.txt alla radice del dominio, dentro Agentic Browsing, la categoria sperimentale dedicata alla navigazione degli agenti: esito positivo se il server lo restituisce, «non applicabile» se risponde con un 404, perché pubblicarlo resta facoltativo, problema segnalato se la richiesta finisce in un errore del server. Il controllo verifica che il file esista, non che cosa contenga.

What shows up in the logs, and how far it goes

Requests to /llms.txt show up in the server log, the record where every access logs an IP address, date, requested resource, and user agent, the string a program uses to declare who it is. Weeks without a single line mean that during that period, on your domain, no program requested the file.

You can’t rely on the user agent alone. It can be faked in an instant, and in the log a scraper posing as a known bot looks identical to the real thing. That’s why operators publish the IP ranges of their own infrastructure: OpenAI for GPTBot, OAI-SearchBot, and ChatGPT-User, Perplexity for PerplexityBot and Perplexity-User. Comparing the source address against the declared range separates legitimate traffic from anything just borrowing the name.

Visits from SEO tools and validators point to someone checking whether the file exists and whether it follows the format. A visit from GPTBot documents a crawler pass meant for training, and how much that download shapes the way a model talks about your brand is something no log can show. ChatGPT-User and Perplexity-User can request pages in response to a user’s action, and when they do, they show up in the log under their own user agent.

A stronger signal shows up when a request for the file is followed shortly after, from the same infrastructure, by requests for the URLs the file lists. It’s still just a signal, because the system’s reasoning isn’t observable: the fetch proves someone downloaded the document, not that they followed the guidance inside it.

What happens after an agent opens the file

When an agent downloads your llms.txt, the proposal expects it to look through the document for useful information and follow the relevant links. Whether it actually uses or cites one of those resources happens later, based on criteria llms.txt doesn’t control.

A generative answer can come from knowledge the model absorbed during training, frozen until the next update, or from information retrieved on the spot – that’s the difference between GEO and AEO. Your llms.txt can be downloaded by crawlers that feed training, and what happens to it afterward isn’t visible from outside. It only enters live search if someone actually asks for it.

For AI Overview and AI Mode, Google describes query fan-out, breaking a question into multiple related queries, and states that supporting pages are identified through Search systems. These generative experiences rely on ranking and quality systems and pull pages from the index, so being indexed and snippet-eligible is the baseline condition for one of your pages to even enter that pool. The full path of an answer has other steps, and llms.txt isn’t among the ones Google names. Other assistants run on different infrastructures, indexes, and agents, so a universal rule built on a single product doesn’t hold up.

A page still needs to be reachable, clear about what it covers, current, and relevant to the request. The title link and snippet are still how your page shows up in search results; structured data still describes entities and properties wherever it’s supported, and Google notes there’s no markup dedicated to visibility in generative answers. A map gets a system to a resource; choosing it as a source is a separate step, with different criteria.

Where an agent actually has a job to do

Software documentation is the case where the proposal claims the heaviest use. A coding agent needs to pull up a function’s signature, check a parameter, read the authentication procedure, understand how a library handles a specific behavior; in a large body of documentation, an llms.txt can cut down on unnecessary fetches. Among the adoption examples, the proposal cites the developer docs for OpenAI, Anthropic, and Gemini, and several documentation platforms generate the file automatically for every project they host.

Outside of software, the proposal gives examples with little in common besides an agent that needs to retrieve information for a clear purpose: a company laying out its structure and policies, a school organizing course information, a personal site someone could use to piece together a resume.

On a blog, an ecommerce site, or a corporate website, we don’t have solid evidence of a direct benefit for visibility in AI answers. Filling out the file still means choosing which resources to make easier for an agent to find and describing what each one does, and on a site that’s already in good shape – valid sitemap, accessible content, clean pages, consistent structured data – that’s the strongest reason we have to publish it.

Even a file generated by a CMS reflects a selection, shaped by the generator’s own rules. The spec lists among its integrations Mintlify and GitBook for documentation, Yoast SEO and AIOSEO for WordPress, Wix for sites built on its platform, and the 2025 Web Almanac finds llms.txt on 2.1% of sites, with 39.6% of those instances generated automatically by a plugin.

A generator lines up the CMS taxonomy, and anyone who’s never opened the file finds things in there they wouldn’t have chosen. Pages show up with whatever title they have in the backend, so a homepage might appear with its actual SEO title or with some working name nobody ever bothered to change. Alongside the content you’ve built your expertise on, you’ll find landing pages for expired promotions, pages made for a single channel or a discount code, flash news, categories, glossary entries, all sitting at the same level. If the plugin also filled in the description line, it pulled it from the site’s description, in whatever language that’s set to. Add it all up and that’s the portrait of your brand you’re handing to whoever reads the file, and nobody actually wrote it.

If the file already exists, open it and check what it’s declaring on your behalf. If it doesn’t exist yet, writing one makes sense mainly when you have a clear agentic use case – documentation, APIs, procedures that coding assistants consult – or when you actually see requests coming in through your logs. On an editorial site without that evidence, llms.txt is a side task compared to the pages you should actually be declaring in it.

Which pages to include in the file

The basic rule for filling out an llms.txt file is: include a resource when you can say what task it helps an agent accomplish. “This page is good” and “this page is recent” aren’t tasks. Multiple pages on the same topic make sense when they serve different functions: API documentation and reference, a procedure and a limits page, a guide and a changelog can all sit in the same file. They overlap when they answer the same request with equivalent content, and that’s where a duplicate entry hands the agent an ambiguity that’s actually yours. It also helps if every page can be understood on its own, outside the navigation path built for human visitors.

On a topic you’ve owned for years, there’s rarely just one page. There’s the guide, the update that followed it, the category that’s grown stronger in the meantime, the commercial landing page that covers the same topic from a different angle. Listing them all without asking what job each one does just avoids the decision. Telling apart pages that coexist from pages competing for the same query requires looking at the site from the outside, with data you can’t eyeball on your own, which is where SEOZoom helps.

When several of your pages compete for the same request

You can see which of your pages overlap on a topic in the project’s Cannibalization report. Each row shows traffic, keyword count, AI mentions, and the number of other URLs involved; opening it shows you which URLs they are, how many keywords they share, and how many mentions each one has. The list doesn’t tell you which page belongs in the file; it shows you the URLs sharing keywords and lets you spot the cases worth checking before you compile it. The mentions column adds which of those pages generative systems already cite.

Attiva una prova di SEOZoom e scopri cosa dice di te l'AI

With Monitored Keywords you check which URL on your domain shows up in the tracked SERP. Pick an entry from the list, and the right-hand panel shows it alongside the page’s and domain’s authority scores. If it’s different from the address you were about to put in the file, you have a decision to make: align with the URL that’s currently ranking in that SERP, or keep yours and know why. When two of your pages compete for the same query, telling them apart or merging them is the work that makes the entry worth including.

Cannibalization and Monitored Keywords are most useful when the resources you’re declaring target organic queries: editorial sites, ecommerce, corporate sites. For technical documentation, the criterion is still the job the page does, and an API reference belongs in the file because it solves a specific problem for an agent, not because it ranks.

Whether the destination holds up the role you’re giving it

You check a destination’s technical soundness with SEO Spider, which returns response code, canonical, robots directives, title, description, and headers for every URL. The report flags destinations that need fixing or checking. A URL that returns an error doesn’t deliver the intended resource; for one that returns a redirect, it’s better to point directly to the final destination. A URL marked noindex, or with a canonical pointing to another address, remains perfectly usable by an agent instead, documentation you don’t want in the index is the typical case, and all it asks is that you confirm it’s a deliberate choice, not a leftover. The crawl starts from one address and follows internal links, so a resource isolated from internal linking won’t show up in the report and needs to be checked separately.

A migration can render an otherwise valid llms.txt obsolete. A page changes address, a category gets closed, two pieces of content get merged, a landing page gets replaced, and the CMS keeps exporting the old taxonomy: the file stays in place and hands the agent a correct-looking map to resources that have since changed role.

The file’s syntax, and the notes that make it work

The spec sets the structure. At the top, an H1 title with the project or site name, the only mandatory part. It can be followed by a summary in blockquote, the quote format that in Markdown opens with the > symbol, with the information needed to understand the rest, then any paragraphs or lists without headings, where you explain how to interpret the resources, and finally sections introduced by an H2 heading, each with a list of links where every entry carries the resource name, the URL, and, after a colon, a note.

# Nome del progetto
 
> Che cosa fa il progetto e per chi, in una o due frasi.
 
Indicazioni su come usare le risorse elencate.
 
## Documentazione
 
- [Guida introduttiva](https://example.com/docs/guida.md): requisiti e primi passi
- [Riferimento tecnico](https://example.com/docs/riferimento.md): parametri e risposte
 
## Optional
 
- [Storico delle versioni](https://example.com/docs/versioni.md)

A title and an address say little about what to expect; the note lets you understand what a resource is for before you open it. A description like “all the information about our product” doesn’t help anyone choose; a note that scopes the content, states the resource’s purpose, or flags a particular condition does the job it’s there for. The spec’s own guidance calls for clear wording and no jargon left unexplained.

The summary at the top of the file needs to match the homepage, the about page, and the Organization structured data. A brand description that’s out of sync with those puts a second version of the same identity into circulation, and that’s the version machines will read.

Before publishing, the spec’s author suggests a test: give an AI agent your llms.txt as its only starting point and ask it to answer a few questions about the site’s content. The agent should be able to follow the links and land on the right resources. If it can’t, or if it picks the wrong page, something needs isolating: the selection and the notes are the first places to look, but the issue could also sit in the page itself, in the internal links, or in how you phrased the request.

Markdown versions and the links that make them findable

The spec recommends aiming for content that’s already readable by models and suggests pairing every useful document with a Markdown copy at the same address, adding .md at the end (page.html.md) or in place of the extension (page.md). An agent that opens that copy gets the text without the markup wrapping the HTML version, and the two address formats coexist because real-world implementations had already gone both ways.

So those copies can also be reached by an agent that arrived without going through the index, version 2 specifies the links that declare them. On the original page, an element link with rel="alternate" e type="text/markdown" indicates where the Markdown copy lives, one with rel="describedby" which llms.txt describes that page. The same links travel in the HTTP Link header of the response, a solution that also works for resources without an HTML page to edit, and that you set up on the server or CDN without touching the site’s code.

Whether it’s worth building them depends on who actually visits your domain, and that calculation has to be made alongside the rest of your work on technical content accessibility: the usefulness is much clearer on documentation consulted by coding agents than on an editorial blog, where every copy is one more piece of content you need to keep in sync.

The pages you declare and the sources that show up in AI answers

With llms.txt you’ve told agents which resources you consider useful for certain tasks. The next step is to look at what happens around those same topic areas once generative answers come into play: which URL from your site shows up, whether a different one emerges instead of the one you chose, and which external sources contribute to the answer.

These are two separate layers. In the file you organize and describe the resources you want to make easier to find; in AI answers, you observe the sources that actually appear in the tracked results. Comparing the two tells you whether the page you selected matches the one the system is using, or whether a different setup is emerging instead.

In SEOZoom you can follow this comparison with AI Prompt Tracker. You monitor the prompts that represent the questions and intents relevant to your brand, and in the Web Pages column you see which URLs from your domain show up in the answers. Opening the detail view also shows you the Sources Used, with the address and the passage the system pulled from it.

If the page that shows up is the one you selected, the two signals line up. If a different URL from your site emerges, you can check whether it answers that question better, covers a different intent, or suggests you should revisit your selection. When an external source shows up, you can open it to see what information it’s providing and how much space it takes up in the answer.

A match between the page listed in the file and the one an assistant cites doesn’t prove one caused the other, though: the system could have reached that resource through Search, its own index, internal linking, or some other source, and that path isn’t observable. You’ve measured an agreement between your selection and what shows up in the answer. You haven’t measured an effect.

If the file already exists on your site, check which resources it lists, how it describes them, and whether the URLs are still the ones you want to make available to agents. If you need to build it from scratch, the priority depends on context: documentation, APIs, and procedures consulted by agents give it a concrete function; on a blog, ecommerce site, or corporate site with no signs of actual use, it comes after the work on your pages, their accessibility, and the quality of the information they contain. llms.txt can make a resource easier to find. The value still lies in the resource the agent finds.

The market will not wait.
Take control now.

The only platform to hold your ground on Google and AI engines.

  • Full
    platform (trial included)
  • Strategy demo
    with an expert
  • Support
    in Italian