NVIDIA · unified memory
NVIDIA DGX Spark
The NVIDIA DGX Spark has 128GB of unified memory, roughly 115GB of it GPU-addressable, with 273 GB/s of bandwidth. The largest model that loads cleanly at a quality quant is DeepSeek-V4-Flash (llama.cpp, 8K context).
Unified memory: 128GB shared between CPU and GPU; we count ~90% as GPU-addressable for model weights. This class fits huge models but is bandwidth-bound: a box that fits a 120B-class model can still decode slowly, while prompt processing may land a higher tier — one headline number would mislead in both directions, so speed here is a tier, labeled beta until calibrated.
- Unified memory
- 128 GB
- GPU-addressable
- ~115 GB (90%)
- Memory bandwidth
- 273 GB/s
- Backend
- cuda
- Price (as configured)
- $4,699
Specs as of 2026-07-11 · source
Affiliate link — we may earn a commission. Verdicts are computed before any link is attached.
What an NVIDIA DGX Spark can run
Every model we track, graded at its best clean-fit quant.
Compare NVIDIA DGX Spark
- vs Mac Studio M3 Ultra 96GB
- vs GMKtec EVO-X2 Ryzen AI Max+ 395 128GB
- vs NVIDIA GeForce RTX 4090
- vs Intel Arc B580
- vs Mac mini M4 24GB
- vs NVIDIA GeForce RTX 3060 12 GB
- vs NVIDIA GeForce RTX 3090
- vs NVIDIA GeForce RTX 4060
- vs NVIDIA GeForce RTX 4070
- vs NVIDIA GeForce RTX 4080
- vs NVIDIA GeForce RTX 5090
- vs AMD Radeon RX 7900 XTX
Have a specific model in mind? Check it against the NVIDIA DGX Spark at your own quant and context length.
Check with the NVIDIA DGX Spark preselected →Verdicts computed from published specs and measured GGUF sizes — how the math works.