TokenAssemble

Can an NVIDIA DGX Spark run Qwen3-32B?

Yes. Qwen3-32B needs 37.5GB at Q8_0 with llama.cpp and 8K context; the NVIDIA DGX Spark has 128GB of unified memory (~90% GPU-addressable) (114.7GB usable) — grade A, with decode speed in the painful tier.

NVIDIA DGX Spark · Qwen3-32B · Q8_0

Fits — grade A

PainfulbetaNew hardware class — tier not yet calibrated
weights 34.8 GBKV cache 2.15 GBoverhead 0.52 GBpool 114.7 GBneeds 37.5 GB

assumes llama.cpp · batch 1 · 8K context · fp16 KV cache

Also runs with: Ollama (one-line install) · LM Studio (GUI) · KoboldCpp (GUI)

Rent a GPU that fits

It fits, but only in the Painful tier — usable for a batch job, rough for anything interactive. If you need it responsive, renting a larger GPU by the hour is worth pricing against an upgrade.

Browse GPUs on Vast.ai

Referral link — we may earn a commission. This never changes a verdict: the fit above is computed from the data before any link is attached.

$4,699 MSRP · as of 2026-07-11

Affiliate link — we may earn a commission. Verdicts are computed before any link is attached.

Every quant, graded

NVIDIA DGX Spark × Qwen3-32B, llama.cpp, 8K context.

QuantFile sizeFitSpeed tier
Q8_034.8 GBFits · Apainful · beta
Q6_K26.9 GBFits · Apainful · beta
Q5_K_M23.2 GBFits · Apainful · beta
Q4_K_M19.8 GBFits · Ausable · beta

Qwen3-32B on similar hardware

More models on the NVIDIA DGX Spark

Verdict computed from published specs and measured GGUF file sizes — see how the math works.