Kimi and GLM changed the map. Price the hardware behind them.
The current open-model conversation is sparse, agentic, multimodal, and long-context. Compare Kimi, GLM, Qwen, MiniMax, StepFun, DeepSeek, and gpt-oss by what actually shapes deployment: resident weights, active parameters, context, VRAM, and live GPU cost.
Three lanes matter now
The useful split is no longer small versus large. It is practical sparse models, frontier agentic clusters, and differentiated challengers.
One-card MoE
GLM-4.7-Flash, Qwen3.6, and Kimi Linear put 3B active paths inside much larger models. This is where most teams should begin.
Agentic clusters
Kimi K2.7 Code and GLM-5.2 target long-horizon engineering. Their active paths are efficient; their resident weights still demand serious clusters.
Different bets
MiniMax M3, Step 3.7 Flash, DeepSeek-V3.2, and gpt-oss each trade license, context, runtime maturity, and memory differently.
Start with the workload, then check the resident weights
Filter by agentic coding, reasoning, long context, or multimodal work. Active parameters hint at token-time compute; the VRAM baseline reflects the much larger set of weights that must stay resident.
Deployment patterns for the models that still fit a bursty workload
The practical flow is usually: pick a model on Hugging Face, use a vLLM-compatible runtime, cache weights aggressively, and choose how much platform management you want to own.
Move from model choice to provider and GPU
Compare the lowest tracked vLLM entry points, provider tradeoffs, and model-specific GPU guides without leaving the pricing data.
Compare provider-specific prices for common LLM workloads
These pages start from a provider and narrow its live GPU catalog to vLLM serving, batch inference, fine-tuning, or training-friendly rows.
Model hosting cost by provider
These pages answer the next question after model selection: what does this model cost on a specific GPU cloud?
What the leading open-model labs are publishing now
This feed follows Moonshot AI, Z.ai, Qwen, MiniMax, StepFun, DeepSeek, OpenAI, and Mistral, then screens out merges, repacks, and low-signal variants. We still show release-kind labels because upstream orgs sometimes publish multiple packaging variants around the same core release.
Quick read on what it takes to host each model
This table is optimized for planning: total versus active parameters, context, memory floor, release timing, and a live cost estimate.
| Model | Best For | Total / Active | Context | Minimum Setup | Released | Cheapest Tracked Hosting |
|---|---|---|---|---|---|---|
|
Loading model table...
| ||||||
How the hosting estimates are computed
The catalog mixes curated model metadata with live GPU pricing. Use it as a planning tool, then add headroom for your workload.