WoluTools

Websites & SEO

Find sitemap URLs that your robots.txt blocks

Paste your robots.txt rules and one sitemap. You get each sitemap URL that a Disallow rule blocks, and can check status codes and canonicals for up to 20 live URLs.

or upload a file

or drop it here

TXT, XML · Up to 10,000 urls
Free account needed for your 3 free jobs a day

No file at hand? See the prepared example

Standard toolIncluded · 3 free jobs a dayFree: 3 jobs a dayPro: up to 200 jobs a day · €12.99/month or €89.99/year
  • URL / TXT / XML
  • Up to 10,000 urls
Prepared result preview · fictional sample dataSynthetic robots and sitemap pair
Source
One sitemap URL sits below a disallowed path
Prioritized findings
  • Disallow rules are parsed by line
  • Sitemap locations are read from loc elements
  • Crawlability is not presented as index status

result.json · findings.csv · http-evidence.csv · read-me.txt

BringURL, TXT, XML
GetRule conflicts, separate coverage counts and optional HTTP/canonical observations
PrivacyEncrypted source · 24-hour result

One clear job, from source to download

  1. 1

    Add the source

    Supported formats and limits are visible before the upload.

  2. 2

    Confirm the settings

    Review the exact source, options, units and access before processing.

  3. 3

    Inspect and download

    Check the preview and warnings, then unlock the complete package.

Checking your sitemap against your robots.txt rules

Your rules and one sitemap go in

You paste your robots.txt rules and one sitemap, or upload them as TXT or XML. A sitemap can hold up to 10,000 URLs, and the whole input can be at most 2,000,000 bytes. Sitemap index files are rejected, so pick one child sitemap and run it on its own. If your robots.txt is empty on purpose, tick the box that confirms it. Missing rules are rejected rather than assumed.

The blocked URLs and the optional live check

You get every sitemap URL that a Disallow rule blocks for the user agent you chose. This comparison runs offline on the text you supplied. If you switch on the HTTP check, the tool also looks at status codes and canonical tags in the static HTML for up to 20 sampled public URLs. Offline and live results are counted separately. URLs that were not checked are marked as unverified.

User agent, sample size and what the report cannot say

You choose Googlebot, Bingbot or a generic agent, and you set how many URLs the offline rule check samples. The HTTP check is off by default. It does not crawl your whole site or fetch canonical targets, scripts or page assets. A clean report means the sitemap and the rules agree. It does not predict whether a search engine will index your pages, because that decision is theirs.

Questions before you run it

Can I submit a sitemap index file?

No. A sitemap index is rejected, so pick one child urlset sitemap and run that on its own.

Does the checker fetch my pages?

Not by default. The rule comparison runs offline against the text you supply. The optional HTTP check looks at status codes and static HTML canonical declarations for at most 20 sampled public URLs.

Which crawler rules are applied?

Whichever user agent you select: Googlebot, Bingbot or a generic agent. Disallow rules are parsed line by line and matched against the loc values found in the sitemap.

Does a clean report mean my pages will be indexed?

No. The report describes crawlability against the rules you supplied. Indexing is a search engine decision and is never claimed here, and unchecked URLs stay unverified.

What if my robots.txt is deliberately empty?

Tick the box that confirms it. Missing rules are rejected rather than assumed, because an intentionally empty rule set and a forgotten paste produce very different findings.