Skip to content
nikcli/inference
OpenAI-compatible inference gateway

One API.
Every open model.

Route every request to the cheapest healthy upstream. Built-in caching, request coalescing, and reasoning support — drop-in for Cursor, the OpenAI SDK, or anything that speaks OpenAI.

14+
Providers
40+
Models
Free
100K tokens/mo
$
curl https://inference.nikcli-ai.dev/v1/chat/completions \ -H "Authorization: Bearer $NIKCLI_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"nikseek","messages":[{"role":"user","content":"Hello"}]}'
POST /v1/chat/completions
Auth Bearer key generated in the dashboard
Metered usage events appear on your Usage page
OpenAI-compatible · dashboard keys · usage metering
Capabilities

Route once. Route cheapest.

Drop-in replacement for OpenAI, Anthropic, or any OpenAI-compatible client. One key, every model.

14+ Providers, One Endpoint

Together AI, Fireworks, DeepInfra, Groq, Cerebras, SambaNova, Hyperbolic, Nebius, OpenRouter, DeepSeek, Mistral, Moonshot, Zhipu, and more.

Cheapest Route + Cache

Each call picks the cheapest healthy upstream. Deterministic responses are cached cross-region. Concurrent identical requests coalesce.

Native Reasoning Support

R1 / QwQ + :thinking variants for DeepSeek V4, Kimi K2.6, GLM 5.1, Qwen 3.5, MiniMax — available on demand.

OpenAI SDK Compatible

Drop-in replacement for Cursor, the official OpenAI SDK, or any HTTP client. Just swap the base URL.

Free Tier Included

100K free tokens per month on signup. No credit card required. Upgrade to pro for higher rate limits.

Transparent Billing

Per-call billing at cost-plus margins. You pay what we pay upstream — no markup. Detailed usage logs included.