robots.txt Tester (incl. AI Crawlers)
See which search and AI crawlers your robots.txt allows, rule by rule.
| Crawler | Result | Deciding rule | Group |
|---|
How to use the robots.txt Tester (incl. AI Crawlers)
- Paste the contents of your robots.txt, or load the example.
- Type the URL or path you want to check, such as /blog/post or a full address.
- Read the table: each crawler shows Allowed or Blocked and the exact line that decided it.
- Fix any warnings, then add your own user-agent at the bottom to test a crawler that is not listed.
About this tool
A robots.txt file tells crawlers which parts of a site they may fetch. It is easy to get subtly wrong: a missing slash, a rule placed above the first User-agent line, or a crawler-specific group that silently overrides your general rules. This tester reads your file the way RFC 9309 describes and shows, for one URL at a time, what each well-known crawler would do and which line made the decision.
The list covers the main search engines plus the AI crawlers many site owners now want to control, such as GPTBot, ClaudeBot, PerplexityBot, CCBot, Bytespider and Google-Extended. Matching follows the standard: a crawler uses the group that names it (all such groups are merged), otherwise the * group; the longest matching path wins; Allow wins a tie; * matches any characters and $ anchors the end; and percent-encoded characters are compared in normalised form.
Remember that robots.txt is a request, not a lock. Well-behaved crawlers follow it, but it does not hide pages from people or stop indexing of URLs that are linked elsewhere. Use noindex or authentication for that.
Frequently asked questions
Why is Googlebot allowed when my * group blocks the page?
A crawler only follows the most specific group that names it. If you have a "User-agent: Googlebot" group, Googlebot ignores the * group completely, so you need to repeat any shared rules in that group.
Does blocking Google-Extended remove my site from Google Search?
No. Google-Extended is a control token, not a separate crawler. Blocking it tells Google not to use your content for training its Gemini models, while Googlebot keeps crawling for Search.
Which rule wins when Allow and Disallow both match?
The rule with the longest path wins, because it is the most specific. If both are the same length, Allow wins. The table shows the winning line so you can check this.
Is my robots.txt uploaded?
No. Parsing and matching happen in your browser. The tool does not fetch anything from your site; you paste the file yourself, and nothing is sent anywhere.
Will every AI company respect these rules?
Most large operators say their crawlers follow robots.txt, but it is voluntary. Fetches triggered by a person asking an assistant to open a page may be treated differently from bulk crawling, so check each operator's documentation.