AnaphoraBot
This page exists because our crawler names it. If AnaphoraBot appeared in your server log and you want to know who that is, what it wanted, or how to stop it, everything is below and a person answers the address at the end.
User-Agent: AnaphoraBot/1.0 (+https://anaphorapartners.com/bot; [email protected])
Who runs it
Anaphora, an independent business development and prospect research firm in San Diego, California. We research companies before writing to them, which means reading pages those companies have published.
What it fetches, and what it does not
| Public pages only | Pages a company publishes about itself: leadership and team pages, newsrooms, job posts, and the public records below. Nothing behind a login, nothing behind a paywall, and no attempt to work around either. | Scope |
|---|---|---|
| One page at a time | Requests to the same host are spaced out, at least two seconds apart by default and further apart where the host asks for it or publishes a rate. We are one crawler reading a handful of pages, not a scrape of your whole site. | Rate |
| robots.txt decides | We read robots.txt before the first fetch, apply the rules for AnaphoraBot, and honor Crawl-delay when it is longer than our own spacing. A disallowed path is not fetched, and the refusal is recorded instead. | Permission |
| No forms, no accounts | The crawler reads pages and documented APIs. It does not submit a form on a website, create an account, accept terms on anyone's behalf, or run a site's scripts. Where a site's terms prohibit automated access, no request is issued at all and the refusal itself is what gets recorded. | Behaviour |
| Every fetch is logged | Each request is recorded with its URL, the status we got back, and the time. That record is what lets us answer a question about our own traffic honestly. | Record |
The public records we read
- SEC EDGARFilings, for the companies that file them
- SAM.govFederal registration and award records
- State registriesEntity standing, officers and agents
- The .gov registryThe CISA list, for public sector accounts
- The company itselfLeadership pages, newsrooms, job posts, public permits
How to block it
Add this to your robots.txt and the crawler stops on its next read of the file.
User-agent: AnaphoraBot
Disallow: /
If you would rather not wait, write to us and we will add your host to the list the crawler refuses before it issues a request.
Reach a person
To ask what we hold about you, to have it corrected or deleted, or to stop hearing from us, the same address works and the privacy policy says what happens next.