Qwen · Current all-rounder

Host Qwen3.6 35B-A3B

A current Qwen workhorse for repository-scale coding, reasoning, vision, and multilingual agents.

Customize workload
Parameters
35B
Context
262,144 tokens
Baseline
48GB VRAM
Quantization
Q4 / FP8 inference

Current hosting recommendations

Default interactive workload, 8K context, one concurrent request, always on.

47 qualifying tracked rows
Lowest cost 91.8/100

A40 on RunPod

$0.44/hr$321/mo
Memory
48GB per GPU
Pricing
on-demand
Evidence
75.0/100
Operations fit
85/100

48GB per GPU clears the 48GB planning floor.

Template availability and community or spot pricing can move quickly, so freshness matters.

Production 78.9/100

B200 on GCP

$16.11/hr$11,760/mo
Memory
192GB per GPU
Pricing
on-demand
Evidence
95.0/100
Operations fit
80/100

192GB per GPU clears the 48GB planning floor.

Committed-use SKUs are intentionally excluded; reserved and Dynamic Workload Scheduler SKUs are excluded too. Collection also depends on API-key health, quota, and billing catalog shape changes.

Planning floor

A planning floor derived from catalog VRAM, quantization, context, concurrency, and traffic headroom; benchmark the final runtime before purchase.

Total VRAM
48GB
Per GPU
48GB
GPU count
1

Continue the decision

Change traffic and uptimeRecalculate cost and headroom for your workload. Qwen3.6 35B-A3B on RunPodInspect provider-specific qualifying rows. Compare other modelsReview quality, memory, and hosting envelopes.

Hand this runbook to Claude Code or Codex

Open a terminal in the repository where you want the deployment files, start claude or codex, then paste this prompt. It asks the agent to verify sources and stop before it creates billable infrastructure.

Deployment prompt
Use the infrastructure or model context on this page to create a reproducible open-model deployment. Use this guide as the starting context: https://www.getflops.ai/llms/qwen3.6-35b-a3b. Read the linked model card and provider documentation before choosing hardware or runtime settings. Open every linked primary source and flag any mismatch instead of guessing. Create a deployment folder containing README.md, .env.example with no secrets, a pinned start script or infrastructure manifest, and smoke-test.sh. Make the endpoint OpenAI-compatible where the runtime supports it. Run local/static validation, estimate the billable resources, and stop before provisioning paid infrastructure until I approve.
Guardrails included No secrets in files · verify primary docs · approval before spend