About our crawler
We operate an automated crawler that reads publicly published job listings and public government records. This page explains exactly what it does, how to identify it, and how to contact us or ask us to stop.
If you administer a site our crawler visits and something we are doing is causing you a problem, email crawler@hiringrecord.com. We read every message and we respond. In almost every case we can solve the problem by slowing down rather than by you having to block us.
How to identify our crawler
User agent
HiringRecordBot/1.0 (+https://hiringrecord.com/crawler; crawler@hiringrecord.com)
Source addresses
| Address | Region | Reverse DNS |
|---|---|---|
| 155.138.160.201 | Chicago, US | crawl-01.hiringrecord.com |
Every address we crawl from has a reverse DNS record resolving to hiringrecord.com, and the forward DNS confirms it. If a request claims to be us but does not verify in both directions, it is not us. We keep this table current, and anything not listed here is either stale or an impersonator. We would like to know about it.
We do not use residential proxy networks, rotating consumer IP pools, or any other infrastructure designed to disguise the origin of our traffic. This is a deliberate and permanent commitment, not a description of current practice.
What we collect
- Job listings from applicant tracking system endpoints that vendors publish for syndication
JobPostingstructured data that sites publish in their own page source for search engines- Public filings from federal and state agencies
We record when a listing appears, when its content changes, and when it is removed. That history is the basis of our research.
Hosts we never fetch
Our crawler does not fetch LinkedIn, Indeed, ZipRecruiter, Talent.com, Jobright or Adzuna. Not slowly, not at a low rate, and not from a different address. There is no crawl target for any of those hosts and no adapter that resolves one, so there is nothing in our system that could fetch them by mistake or by a configuration change.
If you administer one of those sites and you are looking at a log line claiming to be HiringRecordBot, it is not us. We would like to know: email crawler@hiringrecord.com and we will confirm in writing that the address is not ours. Every address we crawl from is in the table above and verifies in both directions.
We do hold a small number of rows describing listings on one of those boards. They did not come from a fetch. They came from people using our browser extension who switched contribution on, about pages they had opened themselves — a title, an employer as displayed, a location, the board’s own id for the listing, and the date shown. Never the description, never a page of search results, and never what was searched for. Those rows are marked as member-contributed permanently and are held at our lowest source authority. How that works, and its limits.
What we do not collect
- Anything behind a login. We do not create accounts, submit applications, use credentials, or access any page that requires authentication.
- Personal information. We do not store job descriptions. What we keep from a listing is a fixed list of structured fields — title, department, team, location, work arrangement, employment type, compensation range, posted date, apply URL, requisition ID and a few more — and a cryptographic hash of the description, so we can tell when it changed without keeping what it said. There is no field in our schema for description text, contact details, or the name of a recruiter or hiring manager, so there is nothing to discard: those values have nowhere to be written. The limit of that is the job title, which we do store, and a title naming a person would be stored with it. We have never seen one. If you find one on your record, tell us and we will remove it.
- Applicant or candidate data of any kind.
- Anything a site’s robots.txt disallows for our user agent — on the discovery crawl. See robots.txt below for where that check runs today and where it does not.
How our crawler behaves
- Scheduled once a day, at 04:00 UTC — which is midnight or 1am US Eastern depending on the season. The window is fixed in UTC, not in Eastern time.
- It does not always hold, and here is the record. Over the twenty days from August 23 to September 11, 2026, the crawl ran on eighteen. It started within five minutes of 04:00 UTC on sixteen of those, started late on one after an overnight failure, ran a second time the same day on four when a run was interrupted and had to finish the boards it had not reached, and did not run at all on 9 or 10 September. A second run is the remainder of an interrupted one, not a second pass over boards already read.
- No more than one request at a time to any single host, with a delay and randomised spacing between them
- Conditional requests — not running, and not yet written. Our published data contract says that where a host supports them we send
If-Modified-SinceandIf-None-Matchand honour304 Not Modified. We have not been doing that. The capability was built into the fetcher and no caller ever passed it a validator, so the number of conditional requests we have sent is zero rather than a small share, and a board that has not changed since yesterday is fetched in full anyway. Measured on : of the 264,819 requests we made in the preceding thirty days, 28,206 were answered with a validator we then discarded. That is waste on our side and load on yours, and we are sorry for it. We tried to fix it on September 17, 2026 and stopped, because the obvious fix would have been worse for you than the waste. A304carries no body. Our readers currently treat a response with no body as a listing that is no longer there — so switching the headers on would have made every job that had simply not changed look as though it had been withdrawn, and a posting that looks withdrawn twice in a row is recorded as closed. That would have put fabricated closures on employer records, which is the one error this project exists not to make. The sequence is therefore: teach the readers to tell “unchanged, still listed” from “gone”, and only then send the headers. We are not giving you a date. Until this bullet says the headers are going out, they are not. Workday boards will remain an exception permanently — we read those overPOST, and conditional requests do not apply to aPOST. - Backoff. We honour
Retry-Afterand back off exponentially on429and5xxresponses - Spread out. We distribute requests across the crawl window rather than issuing them in a burst
If our traffic is still too much for you, tell us a number and we will crawl at that rate.
robots.txt
Where we read it. Our discovery crawl — the pass that visits employer websites to find which applicant tracking system they use — fetches robots.txt before every request and checks it per URL, not per host. The rules are applied under RFC 9309: a group naming HiringRecordBot wins over *, the longest matching rule wins, and we fail closed if the file cannot be read at all.
It has run once. That pass ran on and covered 5,332 hosts, and it has not run in full since — so the sentence above describes how it behaves when it runs, not something happening weekly. We are saying so because this page previously reported the 5,332 figure as its “last run” without saying when that was, which reads as a pass that recurs.
Where we do not, yet. The nightly job-board crawl — the pass that reads the applicant tracking systems themselves, which is most of our traffic — does not currently fetch robots.txt. It reads a fixed list of boards we have enrolled, and it has never consulted a robots file for any of them. We are wiring the same check into that path.
What we did in the meantime. On September 11, 2026 we read robots.txt for all 324 hosts that crawl touches, under the user agent above, and checked all 6,646 distinct URLs we have requested against them. None of them disallows any path we fetch. One host serves no robots.txt at all — it answers 401 — which under RFC 9309 §2.3.1.3 is not a refusal.
We are telling you this rather than waiting until the check is wired in, because “we honour robots.txt” was on this page while one of our two crawlers did not consult it, and you would have been right to discount the rest of the page for it.
We hold no cached verdict between runs. A policy is read fresh and kept for the life of one crawl, so a rule you change today applies to us tomorrow.
To set a specific rule for our crawler:
User-agent: HiringRecordBot Crawl-delay: 10 Disallow: /some-path
To block us entirely:
User-agent: HiringRecordBot Disallow: /
We also honour a blanket User-agent: * disallow.
Opting out, and what it means
Email is the channel that takes effect immediately: crawler@hiringrecord.com. We record the block as an enforceable exception that every pass checks before it requests anything, we confirm it in writing, and it applies from the next crawl.
robots.txt blocks our discovery crawl today. That is the pass that visits employer websites to work out which applicant tracking system they use, and it reads robots.txt before every request. Our nightly job-board crawl does not consult robots.txt yet. That is the pass that reads applicant tracking systems we have already enrolled, and we are wiring the same check into it. Until this page says that pass reads robots.txt, do not rely on a robots.txt rule alone to stop it — email us and we will record the block directly.
We want to be straightforward about what happens next, because being vague about this would be worse than stating it plainly.
If you block our crawler, we stop reading your job listings. We will not route around the block, and we will not use another identity to collect the same data.
We continue to use public government records. Filings with federal and state agencies are public records. Blocking our crawler has no effect on those, and we could not remove them from our research even if we wanted to.
Your company page will note that our listing data is limited at your request. We will not describe this as suspicious or draw any inference from it, because we do not think one is warranted. But we will not present incomplete data as though it were complete, so the page will say plainly that listing coverage is limited because you asked us not to crawl, and it will show the public-record information we still have.
Slowing down is almost always better than blocking, for both of us. Email us before you block and we will very likely be able to accommodate whatever the underlying problem is.
Corrections and right of reply
If we have published something about your company that is factually wrong, email corrections@hiringrecord.com. A person reads every report and we correct confirmed errors. There is no review queue and nothing automated behind that address, so we do not quote a turnaround: a number nobody measures would be worth less to you than this sentence. Corrections are logged publicly at hiringrecord.com/corrections, with what was wrong, how many rows it touched, and what changed so that it would not recur.
Any company may also publish a response on its own page, at no cost, with no account required. We do not edit responses and we do not gate them behind any commercial relationship.
Data retention
We retain listing observation history indefinitely, because the historical record is the entire point of the research. We retain raw fetched content for 30 days and then keep only the extracted structured fields and content hashes.
We do not sell or share raw crawled content.
Contact
| Purpose | Address |
|---|---|
| Crawler issues, rate limiting, blocking | crawler@hiringrecord.com |
| Factual corrections | corrections@hiringrecord.com |
| Right of reply | rightofreply@hiringrecord.com |
| Security disclosure | security@hiringrecord.com |
| Everything else | hello@hiringrecord.com |
We aim to respond to crawler and security email within one business day.