Model provider cost

Host Qwen3.6 35B-A3B on Vast.ai

Qwen3.6 35B-A3B needs 1x 48GB+ GPU, and Vast.ai's current cheapest qualifying row is L40 at $0.58/hr.

1x 48GB+ GPU 35B params Current all-rounder Apache 2.0
Cheapest on this provider
$0.58/hr
L40
Monthly estimate
$420/mo
730 hours at the current median
VRAM baseline
48GB
1x 48GB+ GPU
Qualifying rows
6
Updated Jul 30, 2026

Vast.ai rows that can host Qwen3.6 35B-A3B

The cheapest tracked way to host Qwen3.6 35B-A3B on Vast.ai is L40 at $0.58/hr. The overall tracked market floor is $0.44/hr on RunPod, so Vast.ai is $0.14/hr above the current floor.

Price source Vast.ai marketplace offers Marketplace aggregate; quality-eligible offers feed the summary price.
GPU VRAM Per GPU Estimated hourly Estimated monthly Offers Updated
L40 48GB $0.58/hr $0.58/hr $420/mo 1 Jul 30, 2026
Vast.ai marketplace offers 1 offer · Jul 30, 2026 global aggregate
RTX 6000Ada 48GB $0.68/hr $0.68/hr $498/mo 2 Jul 30, 2026
Vast.ai marketplace offers 2 offers · Jul 30, 2026 global aggregate
H100 SXM 80GB $2.20/hr $2.20/hr $1,607/mo 1 Jul 30, 2026
Vast.ai marketplace offers 1 offer · Jul 30, 2026 global aggregate
H100 NVL 94GB $2.76/hr $2.76/hr $2,013/mo 2 Jul 30, 2026
Vast.ai marketplace offers 2 offers · Jul 30, 2026 global aggregate
H200 141GB $4.01/hr $4.01/hr $2,926/mo 2 Jul 30, 2026
Vast.ai marketplace offers 2 offers · Jul 30, 2026 global aggregate
B200 192GB $6.19/hr $6.19/hr $4,518/mo 7 Jul 30, 2026
Vast.ai marketplace offers 7 offers · Jul 30, 2026 global aggregate

Why this setup does or does not fit

VRAM floor

1x 48GB+ GPU

1x 48GB GPU; 80GB for long context or higher concurrency. Long prompts, batching, and KV cache can require extra headroom.

Model quality

Current all-rounder

One of the most balanced new open-weight choices when code, tools, vision, and multilingual work all matter.

Operational note

Qwen 35B

Native 262K context is useful, but the KV cache can quickly turn a nominal 48GB fit into an 80GB deployment.

Qwen3.6 35B-A3B on Vast.ai FAQ

Can I host Qwen3.6 35B-A3B on Vast.ai?

The cheapest tracked way to host Qwen3.6 35B-A3B on Vast.ai is L40 at $0.58/hr.

What GPU memory does Qwen3.6 35B-A3B need?

Our baseline for Qwen3.6 35B-A3B is 1x 48GB+ GPU. The practical recommendation is 1x 48GB GPU; 80GB for long context or higher concurrency.

Is Vast.ai the cheapest provider for Qwen3.6 35B-A3B?

The overall tracked market floor is $0.44/hr on RunPod, so Vast.ai is $0.14/hr above the current floor.

How fresh is this Vast.ai Qwen3.6 35B-A3B cost page?

This page recalculates from the latest tracked on-demand rows. The freshest qualifying Vast.ai row shown here is from Jul 30, 2026.

Compare this setup

Hand this runbook to Claude Code or Codex

Open a terminal in the repository where you want the deployment files, start claude or codex, then paste this prompt. It asks the agent to verify sources and stop before it creates billable infrastructure.

Deployment prompt
Use the infrastructure or model context on this page to create a reproducible open-model deployment. Use this guide as the starting context: https://www.getflops.ai/llms/qwen3.6-35b-a3b/vast. Read the linked model card and provider documentation before choosing hardware or runtime settings. Open every linked primary source and flag any mismatch instead of guessing. Create a deployment folder containing README.md, .env.example with no secrets, a pinned start script or infrastructure manifest, and smoke-test.sh. Make the endpoint OpenAI-compatible where the runtime supports it. Run local/static validation, estimate the billable resources, and stop before provisioning paid infrastructure until I approve.
Guardrails included No secrets in files · verify primary docs · approval before spend