nikcli Inference
A single OpenAI-compatible HTTP endpoint that:
- Routes each request to the cheapest healthy provider for the chosen model.
- Caches deterministic responses (
temperature=0orseedset) across all your traffic, cross-region. - Coalesces concurrent identical requests into one upstream call.
- Supports reasoning natively (R1, QwQ) and via
:thinkingvariants on hybrid models.
Behind the scenes we currently route across 14 providers: Together, Fireworks, DeepInfra, Groq, Cerebras, SambaNova, Hyperbolic, Nebius, OpenRouter, DeepSeek, Mistral, Moonshot, Zhipu, and a self-hosted vLLM fallback.
Base URL
https://inference.nikcli-ai.dev/v1
Authentication
Authorization: Bearer <your-key> on every request. Create keys at /dashboard/keys.
Keys start with nik_live_. The full plaintext is shown only once at creation — copy it immediately.
Next steps
- Quickstart — your first request
- Cursor setup — drop-in for Cursor
- Models — full catalog + aliases