All work
01Distributed LLM Infrastructure2026

InferGate

A cache that only handles repeats is not enough when fifty identical requests can arrive inside the same second.

FastAPIRedisDockerNginxPrometheus

Context & problem

Every call to an upstream LLM provider costs money and time. A naive cache helps once traffic repeats slowly, but under real concurrency, dozens of identical requests can all miss the cache in the same instant and all reach the provider before the first response lands.

Architecture

Incoming inference request hits the gateway.

The interesting engineering: The cache was the easy part

Exact and semantic caching were the obvious first move, and the latency win was immediate. Under concurrency, though, several requests could still miss at the same moment and all call the provider before the first response reached Redis. The cache was working, but the system was still wasteful.

Single-flight coalescing fixed the stampede. Circuit breakers and ordered failover handled the next question: what happens when the primary provider disappears completely? The benchmark that mattered measured both the happy path and the outage, not just a clean best case.

Result

188ms
p95 latency

down from 394ms, no-cache control

97.9%
provider calls removed

2,000 requests, 50 concurrent users

95%
semantic recall

zero false positives, 330 labeled pairs

What I learned

  • A benchmark that only measures the happy path hides the failure mode that actually matters. The stampede only shows up under real concurrency, not in a single-request test.
  • Coalescing and failover are cheap to add once caching exists, but they are the part that makes the system trustworthy under load rather than just fast in a demo.

Source & demo

No hosted demo yet. The gateway runs as infrastructure, not a UI, so the architecture above and the benchmark write-up in the repo are the best way to see it working.

View on GitHub

Next project

CampusCopilot