Bulk robots.txt Checker
The robots.txt file decides which parts of your site search engines are allowed to crawl, and one wrong line can quietly remove an entire section from the index. This bulk checker fetches /robots.txt for every domain in your list — over HTTPS first, with an HTTP fallback — and reports the response status, the number of redirects on the way, the file size, how many Sitemap directives it declares, and whether it blocks all crawlers with a "User-agent: *" group containing "Disallow: /". A 404 status is highlighted as normal rather than an error: it simply means the site has no robots.txt, which search engines treat as "allow everything". Use the tool after migrations and platform changes to make sure the rules you moved did not accidentally forbid large parts of the site, and to find sitemaps you forgot to submit to Search Console. Up to one hundred domains are checked in parallel, and the finished report can be filtered and exported to CSV.
How it works
- Paste up to 100 items, one per line. Duplicates are removed automatically.
- If a line contains a full URL, only its host name is used.
- The check runs in the background; the results page shows live progress.
- Filter the report by status code and export it to CSV.
Rate limit: 5 jobs per hour from one IP address. Results are stored for 7 days.
FAQ
Is a 404 status for robots.txt an error?
No. A 404 means the site simply has no robots.txt, and search engines treat that as "crawling is fully allowed". The row is marked with a regular 404 status instead of an error.
What does "Blocks all crawlers" mean?
The robots.txt contains a "User-agent: *" group with a "Disallow: /" rule — the whole site is closed to all compliant crawlers, including Google and Bing. Usually this is a misconfiguration left after staging.
Why does the report show both https and http?
The checker fetches https://domain/robots.txt first. If that fails or returns a client/server error, it retries over http://. The "Via" column shows the protocol of the attempt that produced the reported result.
How long are the results stored?
Jobs and their results are kept for 7 days and then deleted automatically. Anyone who has the result link can view the report, so do not submit confidential internal domains.