Foundations

Foundations

Verified, Then Refused

Site operators are writing robots.txt rules against crawlers that were switched off years ago, and against crawlers nobody has ever identified. You could read that as carelessness, in which case there is nothing to learn from it. Read it as improvisation and it tells you a great deal: people building a classification system that no one has issued them, out of the only material lying around, which is whatever name a client announces about itself.
What they are sorting for is purpose. Indexing, training, a live assistant checking a page on someone's behalf, a checkout. And that sets up a failure your retry logic has no way to see: your agent proves exactly what it is, and gets turned away for it.
Verified, Then Refused
Site operators are writing robots.txt rules against crawlers that were switched off years ago, and against crawlers nobody has ever identified. You could read that as carelessness, in which case there is nothing to learn from it. Read it as improvisation and it tells you a great deal: people building a classification system that no one has issued them, out of the only material lying around, which is whatever name a client announces about itself.
What they are sorting for is purpose. Indexing, training, a live assistant checking a page on someone's behalf, a checkout. And that sets up a failure your retry logic has no way to see: your agent proves exactly what it is, and gets turned away for it.

Web Bot Auth Signs the Sender, Not the Errand

For years my entire toolkit for identifying a crawler was a user-agent string, an IP range, and a reverse DNS record that any motivated adult could dress up to look wholesome. Web Bot Auth swaps that guess for a cryptographic check, and a narrower one than most coverage suggests. It tells you a key holder sent something to this host inside a time window. Who the agent is working for, and why, are ruled out of scope on purpose. Here's what actually gets signed, and what comes back to your desk.

Web Bot Auth Signs the Sender, Not the Errand
For years my entire toolkit for identifying a crawler was a user-agent string, an IP range, and a reverse DNS record that any motivated adult could dress up to look wholesome. Web Bot Auth swaps that guess for a cryptographic check, and a narrower one than most coverage suggests. It tells you a key holder sent something to this host inside a time window. Who the agent is working for, and why, are ruled out of scope on purpose. Here's what actually gets signed, and what comes back to your desk.
Recognition Reading









