Websites & SEO
Find sitemap URLs that your robots.txt blocks
Paste your robots.txt rules and one sitemap. You get each sitemap URL that a Disallow rule blocks, and can check status codes and canonicals for up to 20 live URLs.
- URL / TXT / XML
- Up to 10,000 urls
One sitemap URL sits below a disallowed path
- Disallow rules are parsed by line
- Sitemap locations are read from loc elements
- Crawlability is not presented as index status
result.json · findings.csv · http-evidence.csv · read-me.txt
One clear job, from source to download
- 1
Add the source
Supported formats and limits are visible before the upload.
- 2
Confirm the settings
Review the exact source, options, units and access before processing.
- 3
Inspect and download
Check the preview and warnings, then unlock the complete package.
Checking your sitemap against your robots.txt rules
Your rules and one sitemap go in
You paste your robots.txt rules and one sitemap, or upload them as TXT or XML. A sitemap can hold up to 10,000 URLs, and the whole input can be at most 2,000,000 bytes. Sitemap index files are rejected, so pick one child sitemap and run it on its own. If your robots.txt is empty on purpose, tick the box that confirms it. Missing rules are rejected rather than assumed.
The blocked URLs and the optional live check
You get every sitemap URL that a Disallow rule blocks for the user agent you chose. This comparison runs offline on the text you supplied. If you switch on the HTTP check, the tool also looks at status codes and canonical tags in the static HTML for up to 20 sampled public URLs. Offline and live results are counted separately. URLs that were not checked are marked as unverified.
User agent, sample size and what the report cannot say
You choose Googlebot, Bingbot or a generic agent, and you set how many URLs the offline rule check samples. The HTTP check is off by default. It does not crawl your whole site or fetch canonical targets, scripts or page assets. A clean report means the sitemap and the rules agree. It does not predict whether a search engine will index your pages, because that decision is theirs.
Questions before you run it
Can I submit a sitemap index file?
No. A sitemap index is rejected, so pick one child urlset sitemap and run that on its own.
Does the checker fetch my pages?
Not by default. The rule comparison runs offline against the text you supply. The optional HTTP check looks at status codes and static HTML canonical declarations for at most 20 sampled public URLs.
Which crawler rules are applied?
Whichever user agent you select: Googlebot, Bingbot or a generic agent. Disallow rules are parsed line by line and matched against the loc values found in the sitemap.
Does a clean report mean my pages will be indexed?
No. The report describes crawlability against the rules you supplied. Indexing is a search engine decision and is never claimed here, and unchecked URLs stay unverified.
What if my robots.txt is deliberately empty?
Tick the box that confirms it. Missing rules are rejected rather than assumed, because an intentionally empty rule set and a forgotten paste produce very different findings.