Match your workload to current GPU options

Set the serving envelope. The planner checks memory fit, current price evidence, and operational tradeoffs.

Workload

Required fields have practical defaults.

Workload type
Traffic shape

Recommendations

Calculating from current tracked prices.

What the planner considers

Memory fit is a hard requirement. Cost, headroom, evidence confidence, and operational fit determine the ranking.

Hand this runbook to Claude Code or Codex

Open a terminal in the repository where you want the deployment files, start claude or codex, then paste this prompt. It asks the agent to verify sources and stop before it creates billable infrastructure.

Deployment prompt
Use the infrastructure or model context on this page to create a reproducible open-model deployment. Use this guide as the starting context: https://www.getflops.ai/planner?model=glm-4.7-flash. Read the linked model card and provider documentation before choosing hardware or runtime settings. Open every linked primary source and flag any mismatch instead of guessing. Create a deployment folder containing README.md, .env.example with no secrets, a pinned start script or infrastructure manifest, and smoke-test.sh. Make the endpoint OpenAI-compatible where the runtime supports it. Run local/static validation, estimate the billable resources, and stop before provisioning paid infrastructure until I approve.
Guardrails included No secrets in files · verify primary docs · approval before spend