By Lyle Heartman

In May, RubyGems had to shut off new account registrations for four days. If you haven't followed the story: the Ruby package registry got hit by what researchers say were OpenAI agents running a web-lookup task. A new account every two or three minutes, hundreds of packages pushed, most of them just scraped web pages (UK council calendars and the like) instead of code.

They didn't exactly hide it. Files were named hack.rb, evil.rb, exploit.rb. Researchers also say the agents poked at a then-unknown bug that could have exposed users' API keys. RubyGems says it found no sign any keys were actually taken.

OpenAI's explanation was that the agents were using RubyGems as a sort of makeshift browser during a training run where they didn't have open internet access. Maybe that's true. It doesn't help much if you're a volunteer maintainer spending your week cleaning it up.

I'm not writing this to pile on OpenAI. I'm writing it because the thing that happened to RubyGems can happen to a much smaller site, and most small sites are in a worse position to handle it.

The limits you have were designed for people

Most of the protections on a typical site assume a human on the other end. One account per email. A 20MB upload cap. Maybe a CAPTCHA on signup.

An agent swarm doesn't care about any of that. If the cap is 20MB, it splits the file into pieces and spreads them across forty fresh accounts. Researchers say the RubyGems agents got past email verification with throwaway addresses. And a $10 VPS sized for normal traffic will fall over long before anyone notices.

So the question changes. It's not "how fast can my site go." It's "how fast should anything be allowed to go."

Start with the cheap stuff

Block regions you don't serve. If all your customers are in Canada and the US, there's no reason to accept connections from everywhere. It won't stop anyone using residential proxies, but it cuts a lot of noise for basically free.

Slow things down on purpose. We spend years making sites faster. A fast site is also a fast site to scrape. The trick is knowing where "fast" stops being human, and that's something you can measure.

Pull a few weeks of analytics and look at how real people move through your site. Not the average, the extremes. What's the quickest anyone has ever filled out your signup form? The shortest gap between two page views from a real customer? The most pages anyone has looked at in a minute? That fastest real human is your floor. Anything faster than that isn't a person.

Some examples of what that looks like in practice:

  • Signup forms. Even someone with autofill and a password manager usually takes 4 to 5 seconds to get through name, email, and password. A form submitted 400ms after the page loaded wasn't filled in by a person.
  • Browsing. People read, scroll, and click. Your fastest real visitor might hit 15 or 20 pages in a minute when they're skimming product listings. A client pulling 200 pages a minute is doing something else.
  • Checkout. Picking a product, entering shipping, entering a card, confirming. Even a repeat customer with everything saved takes 20 or 30 seconds. Checkout in 3 seconds is a script.
  • Uploads. Someone uploading files has to find them, pick them, maybe rename them. Ten uploads in ten seconds from a brand new account doesn't happen.

Don't just watch for too fast, watch for too even. Humans are messy. They pause for 2 seconds, then 11, then 4. A client hitting your site every 3.0 seconds on the dot, hour after hour, is a script with a sleep timer. Bots also tend to be too thorough, visiting every page in order, never going back, never wandering off to the About page.

Once you have your numbers, set your limits just past the human extreme and add delay only beyond that line. A real person never feels it because they never get there. A bot runs into it constantly. Randomize the delay too (say 800ms to 2 seconds instead of a flat 1 second) so it's harder to measure and tune around.

Your numbers won't match mine. A recipe site and a stock trading dashboard have very different "fastest human" profiles, so measure your own before you set anything.

Throttle progressively. Something like this, per minute:

  • 1st request: no delay
  • 2nd request: 100ms
  • 10th request: 15 seconds

Do this at Nginx, Caddy, or Cloudflare, not in your app. On a small server you don't want throttled connections sitting in your worker slots.

Treat signup and upload as your riskiest pages. Rate limit signups per IP and per network range. Block disposable email domains. Track upload volume per IP or device, not just per account, because that's how you catch the file-splitting trick.

Understand what a CAPTCHA is actually for. Most people think a CAPTCHA proves you're human. It doesn't, not really, and it hasn't for a while. Image models solve those grids, and solver farms will clear them for a fraction of a cent each. If the goal were to perfectly separate people from bots, CAPTCHAs lost that fight years ago.

What they're good at is being annoying. Every challenge costs something: a few seconds, a call to a solver service, a bit of money, a slower pipeline. For one person filling out one form, that's a minor irritation. For a bot trying to create 5,000 accounts, it's 5,000 tolls. The point was never to catch every bot. The point is to make your site more trouble than it's worth, so the bot operator decides the next site over is an easier target.

The catch is that humans pay the toll too. Slowing people down is the side effect, not the goal. So the real question is how to make the toll land mostly on bots. That's where the newer options come in:

  • Invisible scoring (Cloudflare Turnstile, reCAPTCHA v3) checks quietly in the background and only puts a challenge in front of traffic that already looks suspicious. Most real visitors never see anything.
  • Proof of work makes the browser burn a few hundred milliseconds of CPU. One visitor won't notice. A fleet of thousands of bots pays for it in real compute. Anubis is a good ready-made option for AI scrapers.
  • Passkeys for high-value actions. A registered authenticator is much stronger proof than any puzzle.

Keep the visible "click all the traffic lights" CAPTCHA as a last resort step-up, not the front door.

When you need more

Know what you're defending against. OWASP keeps a list of 21 automated threat types. Different pages attract different ones:

PageLikely threatFirst thing to add
LoginCredential stuffingRate limit, breached-password check, MFA
SignupFake accountsEmail verification, velocity limits
Search / catalogScrapingPer-identity rate limits
CheckoutScalping, card testingQueue, purchase limits, 3D Secure
Public APIScraping, scanningKeys, per-key quotas, signed requests
CommentsSpamReputation, delayed publishing

Don't rate limit on IP alone. Residential proxies make that easy to dodge. Layer it: per IP, per session, per logged-in user, per endpoint, and per ASN (datacenter traffic on a consumer site is a flag).

One thing people get wrong on login: use two separate buckets, one keyed on username and one on IP. A single combined ip:username key lets one IP try as many usernames as it wants, which is the whole point of credential stuffing. Use sliding windows or token buckets, and return a plain 429 that doesn't say which limit was hit.

Fingerprint the connection. Bots rotate IPs but usually keep the same client stack. JA4 TLS fingerprints (skip JA3, Chrome's randomized extensions broke it), HTTP/2 settings, and Client Hints that don't match the handshake. If it says Chrome but shakes hands like Python, it's lying. Save canvas and WebGL fingerprinting for high-risk flows. It's invasive.

Set a couple of honeypots. A hidden field humans never see:

<div aria-hidden="true" style="position:absolute;left:-10000px;width:1px;height:1px;overflow:hidden;">
  <label for="company_url">Leave this empty</label>
  <input type="text" id="company_url" name="company_url" tabindex="-1" autocomplete="off" />
</div>

If company_url comes back filled, drop the submission quietly.

And a bait path in robots.txt that nothing on your site links to:

User-agent: *
Disallow: /internal-archive/

Anything that visits it read your robots.txt and went there anyway. Flag it.

Tarpit instead of blocking. A 403 tells the bot exactly what tripped it. Escalate instead: log it, then challenge, then serve slow jittered responses, then disable sensitive actions like checkout. If you're sure, hold the account for review rather than deleting it so you keep the evidence. Same idea as the CAPTCHA: you're not trying to win every round, you're trying to make the whole thing not worth it.

Poison scrapers, carefully. For confirmed scrapers you can serve slightly wrong prices or fake stock counts. Only at very high confidence. Show a real customer a wrong price and you've got a consumer protection problem. Canary listings (fake items that only exist on your site) are safer and give you proof if they show up elsewhere.

Harden login. Check passwords against HaveIBeenPwned (only the first five characters of the hash leave your server). Step up to MFA when patterns look odd. Same error message for wrong username and wrong password.

Log every decision. Timestamp, route, IP, ASN, fingerprint, user agent, hashed user ID, and what you decided and why. You can't tune rules you can't see.

Don't forget the humans

VPN users, privacy browsers, and screen readers can all look a bit like bots. Challenge before you block. Keep an accessible fallback. Hash fingerprints, don't keep raw signals forever, and mention bot protection in your privacy policy.

The mistakes I see most: blocking every unusual user agent, trusting one vendor's black box with no fallback, CAPTCHAs on every login, rules with no logging, and hard blocks on the first signal.

Where this leaves us

Nobody at RubyGems was the target of a planned attack, and it still cost them four days. Big platforms can absorb that. Small ones often can't. Until AI companies are held to real standards for how they test agents on the open internet, it's worth assuming some share of your traffic isn't human and building for it.

Sources