The robots.txt file allows you to provide instructions to search engine crawlers and other automated bots regarding which parts of your website should or should not be crawled.
It is one of the most commonly used files for basic search engine optimization (SEO) and website indexing management.
Table of Contents
What Is the robots.txt File?
The robots.txt file contains rules that are intended exclusively for web crawlers.
Unlike .htaccess, the robots.txt file:
- Does not affect how visitors access your website.
- Does not provide security protection.
- Does not block access to files or directories.
- Only provides guidance to crawlers that choose to respect its rules.
Important
The robots.txt file relies on crawler compliance. While major search engines such as Google and Bing respect its rules, malicious bots may ignore them entirely.
Basic Syntax
Allow a Bot to Crawl a Directory
User-agent: BotName
Allow: /directory/
Prevent a Bot from Crawling a Directory
User-agent: BotName
Disallow: /directory/
Specify a Sitemap
Sitemap: https://yourdomain.com/sitemap.xml
Common Directives
User-agent
Defines which crawler the rule applies to.
Example:
User-agent: Googlebot
Wildcard
Applies the rule to all crawlers:
User-agent: *
Disallow
Prevents crawling of a specific path:
Disallow: /private/
Allow
Explicitly allows crawling of a path:
Allow: /images/
Sitemap
Provides the location of your XML sitemap:
Sitemap: https://yourdomain.com/sitemap.xml
Where Should the robots.txt File Be Located?
The file must be placed in the root directory of your website.
By default, the location is:
public_html/robots.txt
Example directory structure:
public_html
├── robots.txt
├── index.php
├── wp-content
└── wp-admin
Once uploaded, the file should be accessible at:
https://yourdomain.com/robots.txt
Recommended robots.txt for WordPress
A commonly used configuration for WordPress websites is:
User-Agent: *
Allow: /wp-content/uploads/
Disallow: /wp-content/plugins/
Disallow: /wp-admin/
Disallow: /readme.html
Disallow: /refer/
Sitemap: https://yourdomain.com/post-sitemap.xml
Sitemap: https://yourdomain.com/page-sitemap.xml
Replace:
yourdomain.com
with your actual domain name.
Why Include a Sitemap?
A sitemap helps search engines discover and index your content more efficiently.
Typical WordPress sitemaps include:
Posts
https://yourdomain.com/post-sitemap.xml
Pages
https://yourdomain.com/page-sitemap.xml
Many SEO plugins generate these automatically.
Popular options include:
- Yoast SEO
- Rank Math
- All in One SEO
Testing Your robots.txt File
After creating the file:
- Upload it to the website root directory.
- Visit:
https://yourdomain.com/robots.txt
- Verify that the file contents are displayed correctly.
You can also test the file using search engine webmaster tools to confirm that the rules are being interpreted as expected.
Important Limitations
The robots.txt file:
- Does not secure sensitive content.
- Does not prevent direct access to files.
- Does not stop malicious crawlers.
- Should not be used as a security mechanism.
For truly private content, use:
- Password protection.
- Authentication.
- Access restrictions via
.htaccess. - Application-level permissions.
Summary
The robots.txt file allows you to control how search engines and other crawlers interact with your website.
To use it:
- Create a file named:
robots.txt
- Place it in:
public_html
- Add rules using:
User-agentAllowDisallowSitemap
- Verify it is accessible at:
https://yourdomain.com/robots.txt
A properly configured robots.txt file can help improve indexing efficiency and give search engines clearer instructions about which parts of your website should be crawled.