I thought rate limiting was a solved problem. Then my AI agents hit production
Keeping requests to a service under some rate always seemed simple to me, and the whole problem a little overblown. Then I started working closely with AI agents. I've been through the euphoria phase: you connect an age
Keeping requests to a service under some rate always seemed simple to me, and the whole problem a little overblown. Then I started working closely with AI agents.
I've been through the euphoria phase: you connect an agent to your company's data, and it starts producing real insights. The trouble starts when those solutions move to something like production scale. We run heavily on Google Cloud, and suddenly a lot of applications slow down because they hit quotas. Resources get burned at a frightening rate, both on GCP services and on the models behind the agents. Tokens stop being merely expensive; they become worth their weight in gold.
So the main task becomes counting and limiting calls to the services you depend on.
governor: perfect, on one pod
For Rust, there's an obvious answer: governor. It's a great crate, and it does exactly what it says: it limits requests within one process.
The joy fades the moment the service scales horizontally in Kubernetes. Each pod enforces its own limit and knows nothing about the others. Say an API gives you 600 requests a minute, and you set the governor to 10 a second. With eight pods, the fleet sends up to 80 a second. The provider answers with a storm of 429s; the retries pile up, and you pay for calls that do nothing.
Redis: works, but come on
The standard fix is a shared counter in Redis. And it works. But come on — how many extra services do we have to stand up for every little thing? It's one more network hop on every single call, and one more component to deploy, scale, monitor, and get paged for. When Redis is slow, everything behind the limiter is slow. When Redis is down, you choose between no limits and no service.
This problem felt too common for those to be the only two options. So I wrote a crate for it.
What rateguard does
rateguard gives a whole fleet of instances one shared limit per key — with no Redis, no central service, and no network call on the request path.
Every pod embeds a Guard. Each decision is made locally, in about 25 nanoseconds. In the background, the pods talk to each other over UDP and agree on how to split the limit between them.
For outbound calls, the pattern is: ask before you call, and wait if told to.
use rateguard::{Decision, Guard};
let guard = Guard::builder()
.bind("0.0.0.0:7946")
.advertise(pod_ip) // where the other pods reach this one
.seeds(["10.0.0.1:7946", "10.0.0.2:7946"])
.limit(10) // 10 calls/s for the whole fleet
.spawn()?;
// Before every call to the quota-limited API:
loop {
match guard.check("vertex-ai") {
Decision::Allow => break,
Decision::Deny { retry_after } => tokio::time::sleep(retry_after).await,
}
}
call_the_api().await?;
Eight pods or eighty, the fleet stays at 10 calls a second.
How it works, in four lines:
- Shares follow demand. Each pod gossips how many calls it is trying to make, and takes a proportional slice of the limit. A busy pod gets more; an idle one doesn't sit on a share it doesn't use.
- Only busy keys are coordinated. Quiet keys run locally on a small fixed share and generate no network traffic at all.
- Nothing on the call path can fail. If the other pods disappear, each keeps enforcing the last share it knew. What happens during a network split is a setting (HoldDown, Quorum, or Optimistic), with a stated bound for each.
- It's tested like a distributed system. The protocol runs in a deterministic simulator — partitions, 30% packet loss, GC pauses, rolling restarts — and the README carries an accuracy report that CI keeps up to date.
What it doesn't do (yet)
I'd rather you hear the limits from me:
- It counts calls, not tokens. If your LLM quota is in tokens per minute, rateguard limits how often you call, not how much each call costs. Weighted checks are something I'm considering — tell me if you need them.
- One limit per Guard. Every key of a Guard shares the same limit in 0.1; different quotas mean different Guards.
- Seeds are IP addresses for now. Joining through a Kubernetes headless service by DNS name already came in as a pull request from a contributor and is in review.
- Trusted network only. Gossip is neither encrypted nor authenticated; keep the port inside your cluster.
- Not for exact counting. Billing, prepaid credits — anywhere one extra call has a real price — still wants a central counter.
Try it
[dependencies]
rateguard = "0.1"
It's 0.1, and I'd love feedback — above all from anyone fighting the same quota problems with agents, and anyone who finds a failure scenario it doesn't survive.
📦 crates.io/crates/rateguard
💻 github.com/donmiro/rateguard
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.