Skip to content
nikcli/inference

Models

All models are open-weight. Pricing is what you pay (USD per 1M tokens). The gateway pays less upstream — your margin shows up in nikcli.marginUsd.

Aliases

These are the easy-to-remember names. They resolve to the canonical model.

AliasResolves toContextIn $/MOut $/M
nikcli-miniqwen-3.5-flash1M$0.08$0.33
nikcli-fastdeepseek-v4-flash1.05M$0.14$0.28
nikseekdeepseek-v4-pro128K$8$32
nikcli-maxkimi-k2.6256K$8$24
nikcli-reasondeepseek-r1-0528164K$0.63$2.69
nikcli-coderdevstral-2128K$8$24
nikcli-visionllama-4-scout1M$4$12
nikcli-freeminimax-2.5-free205K$0$0

DeepSeek

ModelContextIn $/MOut $/M:thinking
deepseek-v4-pro128K$8$32✅
deepseek-v4-flash1.05M$0.14$0.28✅
deepseek-v4-flash-free1.05M$0$0✅
deepseek-v3.2131K$0.32$0.47✅
deepseek-v3128K$5$15✅
deepseek-r1128K$5$20native
deepseek-r1-0528164K$0.63$2.69native
deepseek-r1-distill-32b128K$4$16native

Kimi / Moonshot

ModelContextIn $/MOut $/M:thinking
kimi-k2.6256K$8$24✅
kimi-k2.5256K$7$21✅

GLM (Zhipu)

ModelContextIn $/MOut $/M:thinking
glm-5.1200K$10$30✅
glm-5200K$8$24✅
glm-5.1-free200K$0$0✅

Qwen

ModelContextIn $/MOut $/M:thinking
qwen-3.5-72b128K$12$36✅
qwen-3.5-32b128K$6$18✅
qwen-3.5-14b128K$3$9—
qwen-3.5-flash1M$0.08$0.33✅
qwen-3.6-max256K$1.30$7.80✅
qwq-32b128K$8$24native

Llama (Meta)

ModelContextIn $/MOut $/M
llama-4-scout1M$4$12
llama-4-maverick1M$6$18
llama-3.3-70b128K$59$79

Mistral

ModelContextIn $/MOut $/M
mistral-medium-3.5256K$15$45
mistral-small-4128K$6$18
devstral-2 (code)128K$8$24

Other

ModelContextIn $/MOut $/M
gemma-4-31b256K$5$15
gemma-4-26b-a4b256K$4$12
phi-4-mini128K$1$3
phi-4128K$3$9
minimax-m2.7100K$8$24
minimax-m2100K$6$18
minimax-2.5205K$0.19$1.44
minimax-2.5-free205K$0$0

The live catalog is at /v1/models — it includes the upstream provider routes per model and reflects what is actually enabled in production.