Skip to content
nikcli/inference

Caching & Coalescing

The gateway automatically optimizes your traffic through intelligent caching and request coalescing.

Response caching

Deterministic requests — calls with temperature=0 or an explicit seed — are cached by their request fingerprint.

# Cache hit example
curl .../chat/completions \
  -d '{"model":"nikseek","temperature":0,"messages":[{"role":"user","content":"What is 2+2?"}]}'

The response will include nikcli.cache: "hit" for cached responses:

{
  "choices": [{ "message": { "content": "4" } }],
  "nikcli": {
    "cache": "hit",
    "provider": "openrouter",
    "costUsd": 0
  }
}

Cache behavior

ConditionBehavior
temperature=0Automatic cache lookup
Explicit seedGuaranteed cache hit for same seed
temperature>0No caching (non-deterministic)
DefaultNo caching

Cache TTL

Default cache TTL is 1 hour for hits. You can tune this per-request:

{
  "model": "nikseek",
  "temperature": 0,
  "messages": [...],
  "nikcli": {
    "cache": true,
    "cacheTtlSeconds": 7200
  }
}

Request coalescing

When multiple identical requests arrive simultaneously, the gateway coalesces them into a single upstream call.

Request A ──┐
Request B ──┼──▶ Single upstream call ──▶ Response to A, B, C
Request C ──┘

Coalescing kicks in for requests with identical:

Viewing cache status

Every response includes cache info in nikcli:

ValueMeaning
"miss"Not cached, new upstream call
"hit"Served from cache
"coalesced"Shared upstream call with other requests

Cache metrics

See your cache hit rate in the Usage dashboard.