AI Signal 404
Owner of 1.5M-page site PatronView on a year of fighting scrapers: 214:1 bot-to-human page loads, 35K Claude crawls per referred user, Amazon bot referred none (Nick Gray/PatronView)
PatronView’s owner reports a year-long battle against web scrapers, noting a 214:1 bot-to-human page-load ratio, massive Claude-model crawling per user, and that Amazon’s crawler no longer generates referrals.
The overwhelming bot traffic can inflate bandwidth costs, degrade performance for real users, and skew analytics. Successful blocking of specific bots (e.g., Amazon’s) shows that targeted defenses can work, but the continued high bot ratio and AI-driven crawls indicate that many scrapers still bypass protections, requiring ongoing engineering effort.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Bot requests dwarf human page loads, with over two hundred bots for every human visit.
AI model Claude is responsible for tens of thousands of crawls per referred user, highlighting AI-powered scraping pressure.
Amazon’s crawler no longer contributes referrals, suggesting that at least one major bot has been effectively blocked.
THE READ
What the cluster adds up to.
PatronView, a site with roughly one and a half million pages, has spent the past year testing a variety of anti-scraping tactics. The owner’s data shows that despite those efforts, bots still dominate traffic, delivering more than two hundred times the number of page loads that genuine users generate. This imbalance signals that any remaining defenses are either insufficiently comprehensive or being outpaced by scraper sophistication.
Among the experiments, some approaches failed to curb the volume of automated requests, while others now appear to succeed in at least limiting certain crawlers. The most concrete evidence of success is the absence of referrals from Amazon’s bot, implying that a rule or filter now blocks that specific user-agent. However, the continued high bot-to-human ratio indicates that many other scrapers remain active.
From an engineering perspective, implementing and maintaining these defenses likely required significant development time, possibly involving custom request-filtering logic, rate-limiting services, or third-party anti-bot platforms. The cost includes not only the labor to design, test, and deploy rules but also the risk of false positives that could inadvertently block legitimate traffic or search-engine indexing, potentially harming SEO and user experience.
The current state still shows a massive number of Claude-model crawls per referred user, meaning that AI-driven scraping continues to bypass existing safeguards. Consequently, the defenses stop working against sophisticated, high-volume AI agents, and further refinement or new detection methods will be needed to reduce that load without degrading service for real users.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗