Signals tracked for Rayobyte
The signal classes we read to decide whether a company is actually collecting the web at scale for Rayobyte, ordered by how close each one sits to the moment somebody chooses where the bandwidth comes from.
How we rank a signal
Two questions decide a signal's value: did somebody inside the company have to write it down, and does it carry a date. A requirement in a hiring specification is written by the engineer who owns the problem and dated by the posting. The same phrase on a marketing page is written by marketing and dated never. That distinction matters more than the keyword itself, so we split the signal by source rather than by string.
Of roughly three thousand four hundred live postings read across one hundred and ten company boards on 16 September 2026, eighty-one matched one of the ten keywords in your brief and ten named the anti-bot problem in the company's own words. Three of those ten were at companies that sell scraping infrastructure themselves.
One class was tested and demoted during this run. We had been treating a crawler being named and disallowed in a major publisher's robots file as evidence of scale and conflict. Parsing thirty-four of those files showed the named bots arrive in long alphabetical runs drawn from the same public registry: ninety-one percent of wired.com's registry-matching entries sit in a single sorted block, sixty-three percent of nytimes.com's. Publishers are pasting the list wholesale. So the signal proves a user-agent string reached a GitHub repository, not that a publisher noticed that crawler. For a bot that publishes a policy and obeys the exclusion, it proves the opposite of a buying need. Two accounts that had ranked on that basis were cut.
Closest to the decision
Weeks, while the team is being staffedA requirement somebody inside the company wrote down and a recruiter published with a date on it. Seniority and location reveal the budget, and a hire means the problem survived being ignored.
The access problem named in a live hiring specification
Residential proxies, proxy rotation, anti-bot evasion, fingerprint spoofing, TLS or JA3, CAPTCHA handling, stealth orchestration. Read from the Ashby, Greenhouse, Lever, Workable, SmartRecruiters and Recruitee public APIs, so a posting counts only while the board still returns it.
From this screenYipitData: a role first published 25 August 2026 requires deploying both residential and data centre proxies.
The blocked workload named beside the technique
Search results extraction, marketplace and retail coverage, travel fares, ticketing, social platforms. These are the targets that fingerprint by network rather than by request rate, which is what separates a residential conversation from a plain bandwidth one.
From this screenZoomInfo: one role names proxy rotation and anti-bot evasion beside knowledge of search results page extraction.
A crawl team being formed rather than maintained
Two or more crawl roles opened inside a quarter, or a lead role whose remit is to grow the team. Tooling and vendor choices get made while a team is being staffed, not after it is full.
From this screenYipitData: a lead role opened 28 August 2026 to grow a team of web scraping engineers.
Behaviour the crawl leaves behind
Months, and dated by the record itselfA crawler at scale cannot stay invisible. It gets named and catalogued, and that carries a date that does not depend on the company announcing anything. Read for existence and timing, not for volume.
First appearance in a public crawler registry
The community AI-crawler registry carries one hundred and seventy-five declared agents. Its commit history gives the date each one was first catalogued, which dates a crawler's public appearance. It also records whether the operator publishes a crawl policy at all.
From this screenReflection AI: the registry recorded its crawler on 22 August 2026 and describes it as undocumented.
A crawler's own declared address range
Operators that publish a static source range and ask to be allowlisted are disqualifying themselves, not qualifying. We read the declared range to remove accounts rather than to add them.
From this screenParallel: publishes ten static addresses and asks webmasters to permit connections from its designated ranges, so rotation would break the allowlist it wants.
Stated crawl manners
An operator that documents honouring crawl-delay and robots directives is telling you it self-throttles. A crawler that slows down on request is never banned by address, so it has no residential requirement at any price.
From this screenImageSift: states that robots directives targeting its crawler are respected, and honours crawl-delay.
Where the account gets disqualified
Applied before anything is written, not afterMost of the work in a screen like this is removal. These are the checks that cut accounts which pass every keyword filter.
Competitor and wholesale separation
Companies that resell scraping or proxies surface strongly on every keyword, because they write about the problem more clearly than anyone. Supplying them wholesale may be the most rational motion at this price, but it is a different message and a different decision, so they are surfaced separately rather than ranked.
From this screenFirecrawl: describes keeping a proxy pool healthy and beating sites that fight back, and resells scraping as an API.
Bandwidth that is not the company's to buy
Distributed and incentivised crawl networks push the bandwidth cost onto node operators, so there is no central budget to sell into however large the crawl looks.
From this screenTimpi: running a collector node requires registering a Node Access NFT, so the crawl runs on operators' own machines.
Live-entity check
Every operator's own domain is opened on the research date. A crawler can be catalogued and excluded everywhere while the company behind it no longer answers.
From this screenTimpi: timpi.io did not load on three attempts on 16 September 2026.
What each account arrives with
- The company and the specific signal that surfaced it, linked to the primary source with the date that source was read.
- An explicit boundary on every card saying what the evidence does not prove, so a claim never travels further than its source.
- The screened-and-rejected list with the reason each account failed, including the ones cut for being competitors, for publishing a static address range, or for having no dated window.
Accounts are checked against the seller's own customer and churned lists before anything sends. This screen cannot see them.
Signal classes describe categories of public record and the cadence we read them on. They rank research priority and do not prove purchase intent, budget, or an active evaluation.