ZZenvuk
Technical SEO

Robots.txt Generator

Create clean, compliant robots.txt directives with multi-agent crawler groups, sitemap linking, and blocking safety warnings.

FREE NO LOGIN BROWSER BASED

Configuration Presets

Fast-start templates for standard web architectures.

Crawler Directives (User-Agent Groups)

Control which directories search bots may or may not crawl.

User-agent:

Explicitly permitted paths or subdirectories.

Restricted paths excluded from crawling.

XML Sitemap Directives

Appended to the bottom of robots.txt

Search engines use these directives to discover your site hierarchy without third-party pinging.

robots.txt Live Output

User-agent: *
Allow: /
Disallow: /api/
Disallow: /admin/
Disallow: /*?*

Sitemap: https://zenvuk.com/sitemap.xml

Upload this file to the root of your domain at https://yourdomain.com/robots.txt.

What Is a Robots.txt File and Why Is It Critical?

Direct Definition: A robots.txt file is a plaintext configuration stored at the root directory of a web server implementing the Robots Exclusion Protocol (REP). It instructs automated web crawlers which URL paths they are permitted or prohibited from requesting, preserving server crawl bandwidth and preventing duplicate crawl loops.

Search engine crawlers, archival bots, and automated AI scrapers inspect your server’s robots.txt file before requesting any other resource. Proper configuration ensures search engines focus limited crawling resources on high-value canonical pages rather than administrative backends, checkout funnels, or endless faceted navigation filters.

Crawling vs Indexation: The Most Common SEO Misconception

Crucial Distinction: Disallowing a URL in robots.txt does not prevent it from appearing in search engine results. Robots.txt restricts crawl access, not indexation. If external or internal hyperlinks point to a disallowed URL, Google may still index the URL snippet without crawling page content.

To guarantee that a private, low-value, or duplicate page is completely excluded from search indices:

  • Allow crawling in robots.txt so Googlebot can inspect the HTTP response and HTML payload.
  • Include a <meta name="robots" content="noindex, follow"> tag in the HTML head.
  • Alternatively, emit an X-Robots-Tag: noindex HTTP response header for non-HTML assets like PDF files.

Robots Exclusion Protocol Directives Explained

Formalized under RFC 9309, the Robots Exclusion Protocol supports a standardized vocabulary of directives:

User-agent: [name]

Designates the specific bot or spider to which the subsequent rules apply. An asterisk (*) acts as a universal wildcard matching all crawlers unless a more specific agent block exists.

Disallow: [path]

Specifies a path prefix that the matching crawler must not access. Leaving the path empty (Disallow:) explicitly permits all crawling under that user-agent.

Allow: [path]

Overrides a broader disallow rule for a specific child path. Used when a parent directory is blocked but a specific subfolder or resource must remain crawlable.

Sitemap: [absolute-url]

Declares the full absolute canonical URL of an XML sitemap or sitemap index. Independent of user-agent blocks and readable by all conforming crawlers.

Controlling AI Scrapers, LLM Bots, and Generative Crawlers

Modern websites encounter specialized AI bot crawlers collecting content for foundation model training, retrieval-augmented generation (RAG), and live browsing plugins. You can grant full access to search engine discovery crawlers while selectively gating automated AI extractors:

# Permit standard search indexing
User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

# Restrict LLM training and automated scraping
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

Note: Respect for robots.txt is voluntary. Major AI platforms honor these exclusions, but malicious or unverified scrapers may ignore robots directives. Sensitive content must be protected behind authentication.

Frequently Asked Questions About Robots.txt

Does robots.txt prevent a page from being indexed in Google?

No. Robots.txt only governs crawl access, not indexation. If external links point to a disallowed URL, search engines can still index the URL without fetching its content. To definitively prevent indexation, allow crawling and implement a 'noindex' robots meta tag or X-Robots-Tag HTTP header.

Where must the robots.txt file be uploaded?

The robots.txt file must reside in the exact root directory of your domain: https://example.com/robots.txt. Subdirectory placements such as https://example.com/blog/robots.txt are ignored by standard web crawlers.

Is crawl-delay supported by Googlebot?

No. Googlebot ignores the Crawl-delay directive. Bingbot and Yandex support crawl-delay in seconds. For Googlebot crawl rate management, configure crawl frequency settings within Google Search Console if necessary.

Can I block AI bots without blocking search engine crawlers?

Yes. You can declare specific User-agent blocks for AI crawlers like GPTBot, CCBot, ClaudeBot, and Anthropic-ai with Disallow: / while keeping User-agent: Googlebot and User-agent: Bingbot set to Allow: /.

What is the difference between Allow and Disallow?

Disallow instructs matching crawlers not to request URLs starting with that prefix. Allow overrides a broader disallow rule for specific subdirectories or files. Google and Bing evaluate the most specific matching rule by path character length.

How do wildcard patterns (*) work in robots.txt?

An asterisk (*) represents any sequence of characters in standard robots.txt extensions. For example, Disallow: /*.pdf blocks all URLs ending with .pdf, while Disallow: /search?* blocks internal search parameter URLs.

Related Utilities

Continue your workflow

View all tools