Infrastructure for the new open stack

Price a GLM workload before you rent

Start with 3B-active GLM-4.7-Flash, then compare the hardware behind Kimi, Qwen, MiniMax, StepFun, DeepSeek, and the rest of today's sparse, agentic stack.

  • GLM-4.7-Flash
  • Interactive inference
  • No account or API key
Keep GPU prices on your radar Read the weekly brief Follow via RSS
Live starting point Practical GLM inference $0.44/hr A40 on RunPod Updated Jul 27, 2026
Model
GLM-4.7-Flash
Traffic
Interactive
GPU class
32GB+
Check the full recommendation
Need a different workload? Start from a current price floor.
Auditable rows Every visible price carries source, freshness, region, and offer-count context. Marketplace aggregates, public list prices, and provider API rows are labeled so missing or cheap cells are not over-read.
Loading pricing data...

Hand this runbook to Claude Code or Codex

Open a terminal in the repository where you want the deployment files, start claude or codex, then paste this prompt. It asks the agent to verify sources and stop before it creates billable infrastructure.

Deployment prompt
Use the infrastructure or model context on this page to create a reproducible open-model deployment. Use this guide as the starting context: https://www.getflops.ai/gpu-prices. Read the linked model card and provider documentation before choosing hardware or runtime settings. Open every linked primary source and flag any mismatch instead of guessing. Create a deployment folder containing README.md, .env.example with no secrets, a pinned start script or infrastructure manifest, and smoke-test.sh. Make the endpoint OpenAI-compatible where the runtime supports it. Run local/static validation, estimate the billable resources, and stop before provisioning paid infrastructure until I approve.
Guardrails included No secrets in files · verify primary docs · approval before spend