Model provider cost

Host GLM-4.7-Flash on RunPod

GLM-4.7-Flash needs 1x 32GB+ GPU, and RunPod's current cheapest qualifying row is A40 at $0.44/hr.

1x 32GB+ GPU 31B params Best practical GLM MIT
Cheapest on this provider
$0.44/hr
A40
Monthly estimate
$321/mo
730 hours at the current median
VRAM baseline
32GB
1x 32GB+ GPU
Qualifying rows
8
Updated Jul 31, 2026

RunPod rows that can host GLM-4.7-Flash

The cheapest tracked way to host GLM-4.7-Flash on RunPod is A40 at $0.44/hr. RunPod is currently tied for the cheapest tracked market setup for this model.

Price source RunPod GraphQL API Provider API pricing; availability and spot/community rows can move quickly.
GPU VRAM Per GPU Estimated hourly Estimated monthly Offers Updated
A40 48GB $0.44/hr $0.44/hr $321/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate
A6000 48GB $0.53/hr $0.53/hr $387/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate
RTX 6000Ada 48GB $0.84/hr $0.84/hr $613/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate
L40 48GB $0.91/hr $0.91/hr $661/mo 2 Jul 31, 2026
RunPod GraphQL API 2 offers · Jul 31, 2026 global aggregate
RTX 5090 32GB $0.99/hr $0.99/hr $723/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate
A100 PCIE 80GB $1.39/hr $1.39/hr $1,015/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate
A100 SXM4 80GB $1.49/hr $1.49/hr $1,088/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate
MI300X 192GB $2.39/hr $2.39/hr $1,745/mo 1 Jul 31, 2026
RunPod GraphQL API 1 offer · Jul 31, 2026 global aggregate

Why this setup does or does not fit

VRAM floor

1x 32GB+ GPU

1x 48GB GPU for context and batching headroom. Long prompts, batching, and KV cache can require extra headroom.

Model quality

Best practical GLM

A strong current starting point when you want agentic behavior without moving immediately to a multi-GPU cluster.

Operational note

GLM 31B

The 3B active footprint helps token throughput, but all 31B parameters still need to live in memory.

GLM-4.7-Flash on RunPod FAQ

Can I host GLM-4.7-Flash on RunPod?

The cheapest tracked way to host GLM-4.7-Flash on RunPod is A40 at $0.44/hr.

What GPU memory does GLM-4.7-Flash need?

Our baseline for GLM-4.7-Flash is 1x 32GB+ GPU. The practical recommendation is 1x 48GB GPU for context and batching headroom.

Is RunPod the cheapest provider for GLM-4.7-Flash?

RunPod is currently tied for the cheapest tracked market setup for this model.

How fresh is this RunPod GLM-4.7-Flash cost page?

This page recalculates from the latest tracked on-demand rows. The freshest qualifying RunPod row shown here is from Jul 31, 2026.

Compare this setup

Hand this runbook to Claude Code or Codex

Open a terminal in the repository where you want the deployment files, start claude or codex, then paste this prompt. It asks the agent to verify sources and stop before it creates billable infrastructure.

Deployment prompt
Use the infrastructure or model context on this page to create a reproducible open-model deployment. Use this guide as the starting context: https://www.getflops.ai/llms/glm-4.7-flash/runpod. Read the linked model card and provider documentation before choosing hardware or runtime settings. Open every linked primary source and flag any mismatch instead of guessing. Create a deployment folder containing README.md, .env.example with no secrets, a pinned start script or infrastructure manifest, and smoke-test.sh. Make the endpoint OpenAI-compatible where the runtime supports it. Run local/static validation, estimate the billable resources, and stop before provisioning paid infrastructure until I approve.
Guardrails included No secrets in files · verify primary docs · approval before spend