What it takes to run GLM 5.2.
See where GLM 5.2 sits in live OpenRouter demand, then compare its API price, downloadable weights, memory floor, and cheapest tracked GPU setup against the rest of the market.
Where model demand is concentrated, and where it is moving
Bars show share of all routed tokens. The trace shows each model’s last seven completed days. Movement compares that window with the seven days before it.
The popular model depends on the job
Request share measures how often a workload appears. Token share shows where the longer, heavier work lives. Select a workload to see its largest tasks and their model leaders.
The demand leaders with downloadable weights
This is the verified open-weight subset of the market above. Filter by workload, then compare seven-day demand, API price, resident weights, and the cheapest tracked GPU setup.
Deployment patterns for a list dominated by large models
Most usage leaders are cluster-scale and favor warm capacity. For the smaller entries, the practical flow is still: pick the official weights, use a compatible runtime, cache aggressively, and control cold starts.
Move from model choice to provider and GPU
Compare the lowest tracked vLLM entry points, provider tradeoffs, and model-specific GPU guides without leaving the pricing data.
What the leading open-model labs are publishing now
This feed follows Xiaomi, DeepSeek, Tencent, Z.ai, NVIDIA, MiniMax, StepFun, Moonshot AI, Qwen, OpenAI, and Mistral, then screens out merges, repacks, and low-signal variants. We still show release-kind labels because upstream orgs sometimes publish multiple packaging variants around the same core release.
Quick read on what it takes to host each model
This table is optimized for planning: params, context, memory floor, deployment pattern, and a live cost estimate.
| Model | Best For | Total / Active | Context | Minimum Setup | Released | Cheapest Tracked Hosting |
|---|---|---|---|---|---|---|
|
Loading model table...
| ||||||
How ranking and hosting estimates are computed
The catalog combines an OpenRouter usage snapshot, curated model metadata, and live GPU pricing. Use it as a planning tool, then add headroom for your workload.