SEOpeckBot: Our robots.txt Research Crawler
If you found SEOpeckBot in your server logs, it is a small research script run by SEOpeck. It reads the robots.txt files of popular websites to see how they treat AI crawlers such as GPTBot, ClaudeBot and PerplexityBot. It obeys robots.txt, does not follow links and does not collect content for AI training.
The User-Agent string
SEOpeckBot/1.0 (robots.txt study; +https://seopeck.com/page/robots-txt-study)
What it fetches
We work down the Tranco research list (list 94GG2) in rank order until we have 1,000 websites. Domains are checked in batches of 200, so a few just below the cut-off also get a robots.txt request and, where robots.txt allows it, a homepage request. For each domain, SEOpeckBot sends plain HTTPS GET requests in this order:
/robots.txt, following up to 5 redirects. If there is no valid file, it tries once on thewww.(or bare) version of the domain. If the domain does not connect, or its robots.txt redirects to another site, nothing more is fetched.- Your homepage, only if robots.txt allows SEOpeckBot to fetch
/(through aSEOpeckBotgroup, or else the*group). This checks that the domain serves a website. /llms.txt, only if robots.txt allows SEOpeckBot to fetch it.
If robots.txt answers with 401, 403, 429 or a 5xx error, we skip the domain and fetch nothing else.
How often, and how politely
Each site is checked once per study run. The main crawl runs on Thursday 8 October 2026, after a few test runs earlier in October.
The script makes at most 12 requests at a time across all sites and waits at least 2 seconds between requests to the same host. It gives up after 10 seconds without a connection or 20 seconds in total, and runs from our own machine, not a hosting server.
Why we run it
The study looks at which AI crawlers popular websites block or allow, and at common mistakes such as outdated bot names or lines that parsers ignore. It uses the same parser as our free AI Crawler Checker, which you can use to test your own robots.txt.
What we publish
robots.txt files are public by design. We publish a per-site CSV and a write-up on our blog. Each row has the domain, its Tranco rank, the robots.txt URL and HTTP status, whether each crawler is allowed, blocked or partly blocked, and whether the site has Content-Signal lines or an llms.txt file. We do not republish full robots.txt files or anything from your homepage.
How to opt out
Add this group to your robots.txt:
User-agent: SEOpeckBot
Disallow: /
When SEOpeckBot finds it, it fetches nothing else from the domain and leaves it out of the results. Name SEOpeckBot itself: a User-agent: * rule stops the homepage request, but your robots.txt can still be counted.
You can also contact us with your domain. We will remove its row from the published data, including the dataset on GitHub, and leave it out of future runs.
Contact
Questions about SEOpeckBot or the study: contact us.