All articles
Research
10 min read
What we learned routing 3 billion requests
Patterns, surprises and a few hard lessons from a year of production traffic.
Maya Okafor

Most traffic is easy
Roughly 70% of requests we see can be served by small, fast models without any measurable quality loss. The remaining 30% is where frontier models earn their price — and where routing pays for itself.
The biggest savings came not from cheaper models but from caching: near-duplicate prompts are far more common than anyone expects.
Failover is not optional
Every major provider had at least one multi-hour incident this year. Teams with automatic failover barely noticed. Teams without it wrote postmortems.
Next article