ai.txt, llms.txt, and TDMRep are often mentioned in the same breath as robots.txt, but in practice they serve completely different purposes. TDMRep is a formal mechanism for objecting to text and data mining, in accordance with the European DSM Directive. ai.txt primarily serves as an informal policy governing the use of content by AI systems, while llms.txt is a map of the most important content you want to present to language models.
Table of Contents
TDMRep-A Legal Signal for Text and Data Mining
TDMRep was defined by the W3C as a simple protocol allowing content owners to restrict or permit text and data mining in a machine-readable format. It uses the `tdm-reservation` and `tdm-policy` fields, which can be specified in HTML meta tags, HTTP headers, or via the `tdmrep.json` file in the `.well-known` directory. A tdm-reservation value set to 1 indicates that TDM rights are reserved, while 0 indicates no restrictions.
Unlike robots.txt, TDMRep does not control bots’ physical access to resources. It specifies the rights to use content in text and data mining processes, including for training AI models. It is primarily a legal and licensing signal that can serve as an important argument in discussions with AI solution providers.
ai.txt – AI content usage policy
ai.txt is an evolving, informal standard for a file typically located at /ai.txt. Its purpose is to describe the rules for AI systems’ use of a website’s content. In such a file, you can specify whether you accept the use of content for training models, summarizing, quoting, or generating responses, as well as any restrictions and expectations regarding attribution.
In practice, ai.txt serves as a “consent, not access” layer-it signals consent or lack thereof, rather than acting as a technical block. It does not replace the terms of service or server-level security measures, but it helps provide clear guidelines to entities that wish to use your content in a manner consistent with the owner’s wishes.
llms.txt – a content map for language models
llms.txt is a text or Markdown file typically located at /llms.txt, whose purpose is not to block bots but to curate content. Such a file usually contains the website’s name, a short description, and a list of the most important subpages-such as the main offering, documentation, FAQs, and key articles-that are particularly useful for language models.
You can think of it as a map or guide to the content you want the AI to understand well and, if applicable, cite. llms.txt does not replace sitemap.xml or robots.txt-rather, it serves as a supplement that points models to approved, high-quality sources within your domain.
What do these mechanisms control in practice?
TDMRep governs text and data mining rights-it specifies whether data may be used in TDM processes, including AI model training, but does not physically block access to the site. ai.txt outlines the rules for AI systems’ use of content, with an emphasis on consent for training, summarization, and other forms of processing. llms.txt, on the other hand, controls neither rights nor access-it helps models find and understand the most important content on the site.
All three mechanisms serve more as signals to “well-behaved” systems than as hard security measures. Real control over traffic still lies in robots.txt, server configuration, firewalls, and any WAF solutions or CDN-side filters. Therefore, it’s best to treat TDMRep, ai.txt, and llms.txt as supplements rather than replacements for traditional security and SEO tools.