
What Does an llms.txt File Do?
llms.txt is not a required file, and today no major search engine or AI assistant has stated that it reads it in its production systems. The file itself is a plain text file in the site's root directory that lists the site's important pages in Markdown format. Jeremy Howard proposed the idea, and the goal was simple: instead of stripping out menus, ads, JavaScript and template clutter to understand a site, a language model could go straight to a clean list of the pages that are actually useful.
The format is defined in the specification. The only required element is an H1 heading carrying the name of the project or site. Below it come a blockquote summarizing the site in a sentence or two, then free-form description sections without headings, and finally link lists separated by H2 headings. Each link is written in standard Markdown and can take an optional short note. The second version of the specification simplified the structure somewhat and defined links under an "Optional" heading as secondary resources that can be skipped when a shorter context is needed.
The file doesn't have to sit in the root directory. It can also be published under a subpath. For example, /docs/llms.txt covers only the addresses under that directory, and if more than one file applies, the agent is expected to use the most specific one.
What Does Google Say About llms.txt?
Google's position is clear enough to leave no room for interpretation. Its Search Central documentation states that there's no need to create new machine-readable files, markup or Markdown to appear in Google Search and its generative AI features. The same document names llms.txt specifically and says creating the file will neither help nor hurt a site's visibility or ranking, because Search ignores it.
This isn't a defense added after the fact. Gary Illyes of the Search team said at a Search Central event that Google doesn't support llms.txt and has no plans to, and that normal SEO is enough to appear in AI Overviews. When John Mueller was asked whether the presence of llms.txt in Google's own developer documentation meant an endorsement, his answer was one word: no.
In the same discussion, Mueller pointed to a problem that is less technical but more lasting. In his view, the file doesn't give models a way to tell which site is better, because everyone describes themselves favorably in their own file. That's a problem independent of whether the file gets read. The promotional text a site writes about itself isn't the kind of evidence that can be used as a selection or ranking signal.
Who Actually Reads the File?
"Even OpenAI has an llms.txt file" is the most frequently used argument in this debate, and it looks at the wrong thing. Both OpenAI and Anthropic publish llms.txt files for their own developer documentation. But publishing a file isn't the same as reading one.
The picture becomes clear when you look at the official documents in which both companies describe their crawlers. OpenAI points only to robots.txt as the control mechanism it offers site owners for GPTBot, OAI-SearchBot and ChatGPT-User. In Anthropic's document describing ClaudeBot, Claude-User and Claude-SearchBot, llms.txt isn't mentioned at all; the only reference is again robots.txt. In other words, these companies use the file as a kind of table of contents to offer their own documentation to agents, and nowhere do they say they will read the file on your site.
Server logs support this picture too. In an Ahrefs study that examined log records for 137,000 domains, 97 percent of llms.txt files received no requests at all within a one-month measurement window. Even among the small minority that did get requests, most of the traffic came not from AI engines but from SEO audit tools, technology profiling bots and ordinary web crawlers. GPTBot accounted for a small slice of those requests, and ClaudeBot stayed below one percent.
A second layer of data looks at citations. Across about 300,000 domains, SE Ranking found no meaningful relationship between having an llms.txt file and how often a site was cited in large language model answers. These are observational studies by tool vendors, not proof of causation, and each is limited on its own. But two different methods and two different datasets point in the same direction, and nothing so far shows that the file measurably changes citation behavior.
llms.txt and robots.txt Are Not the Same Thing
This is where the most common confusion lies. llms.txt blocks nothing, allows nothing and directs no crawler. It only describes. robots.txt, by contrast, is an instruction file, and AI providers officially state that they follow its directives. Confusing the two is like confusing a building plan with a door lock.
| File | What it does | Do AI engines follow it? | Effect on AI visibility |
|---|---|---|---|
| robots.txt | Determines which addresses crawlers can access | Major providers officially state that they follow it | Direct and decisive |
| sitemap.xml | Provides a list of addresses to crawl | Supported as a standard by search engines | Indirect, on the discovery side |
| llms.txt | Describes important pages in a simplified form for models | No major provider has stated that it reads it in production systems | No measurable effect has been shown |
The distinction that really gets missed is inside robots.txt. AI companies don't run a single bot; they run different bots for different jobs, and each can be managed separately. On the OpenAI side, GPTBot collects content for model training, OAI-SearchBot makes a page eligible to appear as a source in ChatGPT's search answers, and ChatGPT-User only comes into play when a user provides an address. On the Anthropic side, ClaudeBot and Claude-SearchBot are separated in a similar way. If these names and the differences between them get confusing, our AI and SEO glossary offers a quicker reference.
Here is why the distinction matters in practice. If you block all AI bots wholesale because you want to stop model training, the same move also shuts off your chances of being cited as a source in AI search. A site's AI visibility is far more often decided by this kind of robots.txt call than by a missing llms.txt. Managing crawler access page by page and deciding where crawl budget goes is a separate topic, covered in our guide to robots.txt and crawl budget.
Who Benefits From llms.txt?
Sites that publish technical documentation
This is the scenario the specification was designed for, and here the file does real work. When a developer gives a coding assistant the address of your documentation, the agent can work through a single Markdown index instead of visiting pages one by one and parsing HTML. The mechanism isn't search engine discovery but direct referral. In other words, it isn't a crawler finding the file but an agent being handed the address. For a software company publishing API documentation, SDK references, product manuals or installation guides, this is a reasonable investment, and it's exactly why AI labs publish their own files.
Corporate sites and e-commerce
There's no evidence that a local service site, an e-commerce store or a corporate brochure site has gained anything measurable from llms.txt. For these sites, the file is half an hour of work and low risk. The problem isn't the file itself but the meaning attached to it. The statement "llms.txt is essential for AI visibility" isn't true, and a site owner who believes it ends up thinking they've done the work that actually makes a difference when they haven't. It's not a loss of budget; it's a loss of attention.
The cost of an unmaintained file
The file's only concrete risk is going stale. An llms.txt that lists removed pages, changed addresses or descriptions that no longer apply sends the few agents that read it to the wrong place. If it won't be updated every time the site structure changes, not publishing it at all is the more consistent choice. A file that's published and forgotten is worse than no file.
Where That Effort Is Better Spent
Both Google's own guidance and independent data point to the same place. The path to appearing as a source in AI answers runs not through special files but through the page's accessibility and the content itself. In order:
- The page must be crawlable and indexable. A wrong canonical, an unnecessary noindex or a blocked directory takes even the best content out of play.
- The main content must be readable in the HTML. A page where the answer only appears after JavaScript runs is half invisible for retrieval purposes.
- The robots.txt decision must be made deliberately, bot by bot. Training bots and search bots should be evaluated separately.
- The answer must exist within the page as a passage that makes sense on its own. If a section still makes sense when lifted from the page, it's suitable for citation.
- The content must contain something not found elsewhere. A page with its own data, its own test or its own comparison stands apart from a page that repeats the same information in different words.
- What the brand, product and service are must be clear from the text, with no ambiguity.
None of these items requires a new file format; all of them are classic technical SEO and content work. Done together, they produce the only thing that can actually be measured in AI search visibility. llms.txt doesn't replace any item on this list.
If You're Going to Publish the File Anyway
If you want to create one because the cost is low, go in with the right expectations. Three rules are enough to do it properly. The file should follow the specification, meaning it contains an H1 heading, a short summary blockquote and Markdown link lists grouped under H2s. It should include only the pages that are genuinely representative, not every page. And it should be updated when the site structure changes.
If you don't want to write the file by hand, Seobaz's free llms.txt generator crawls the site and its sitemap and produces a draft file. Review the output instead of uploading it as is, because automatically selected pages aren't always the pages the site actually wants to highlight. The tool's job is to prepare the draft; the decision about which pages make the list is yours.
Once it's live, keep your expectations realistic. The file won't change the site's Google rankings, won't increase the chance of appearing in AI Overviews and most likely won't be requested by any AI engine. Even so, publishing it is a defensible choice, because if the standard is adopted in the future, the file will be ready. What isn't defensible is publishing the file and considering the AI visibility work done.
Frequently Asked Questions
Will creating llms.txt hurt my site's Google rankings?
No. Google's documentation is clear on this: creating the file neither helps nor hurts the site's visibility or ranking, because Search ignores it. The only risk of harm is indirect. If the file contains wrong or dead addresses, the few agents that read it get sent to the wrong place.
Do llms.txt and sitemap.xml do the same job?
No. sitemap.xml gives search engines a complete, machine-readable list of addresses to crawl and is supported as a standard by the major search engines. llms.txt is a curated selection that explains in human language which pages are important and why. One is about coverage, the other is a claim about priority. The sitemap is a supported standard; llms.txt is a proposal that hasn't been adopted yet.
Does llms.txt stop AI bots from using my content for training?
No. llms.txt has no blocking power; it only describes. A site owner who wants to limit training use should use robots.txt, and the bots to block are training crawlers such as GPTBot or ClaudeBot. When you make this decision, keep search-side bots separate; otherwise you also shut off the chance of being cited as a source in AI search.
Do I also need llms-full.txt?
It isn't a file defined by the specification. Some documentation platforms started using this name to offer the full text of their docs in a single file instead of a link list, and the practice spread. So it's not a standard but an established habit. For sites that don't publish documentation, producing a second, much larger version of a file whose readership is already uncertain isn't worth the effort.



