InferGate
A cache that only handles repeats is not enough when fifty identical requests can arrive inside the same second.
Context & problem
Every call to an upstream LLM provider costs money and time. A naive cache helps once traffic repeats slowly, but under real concurrency, dozens of identical requests can all miss the cache in the same instant and all reach the provider before the first response lands.
Architecture
Incoming inference request hits the gateway.
The interesting engineering: The cache was the easy part
Exact and semantic caching were the obvious first move, and the latency win was immediate. Under concurrency, though, several requests could still miss at the same moment and all call the provider before the first response reached Redis. The cache was working, but the system was still wasteful.
Single-flight coalescing fixed the stampede. Circuit breakers and ordered failover handled the next question: what happens when the primary provider disappears completely? The benchmark that mattered measured both the happy path and the outage, not just a clean best case.
Result
down from 394ms, no-cache control
2,000 requests, 50 concurrent users
zero false positives, 330 labeled pairs
What I learned
- A benchmark that only measures the happy path hides the failure mode that actually matters. The stampede only shows up under real concurrency, not in a single-request test.
- Coalescing and failover are cheap to add once caching exists, but they are the part that makes the system trustworthy under load rather than just fast in a demo.
Source & demo
No hosted demo yet. The gateway runs as infrastructure, not a UI, so the architecture above and the benchmark write-up in the repo are the best way to see it working.
View on GitHubNext project
CampusCopilot