Ollama · 32GB VRAM
Best Ollama models for 32GB VRAM
34 models we track load cleanly on 32GB — quality quantization, no CPU offload, 8K context. The largest is Qwen3.6-35B-A3B at UD-Q5_K_XL. Verdicts are computed on a NVIDIA GeForce RTX 5090, a real card we have sourced specs for — not a generic 32GB abstraction.
Cards with 32GB: NVIDIA GeForce RTX 5090 · NVIDIA RTX 5000 Ada Generation.
| # | Model | Params | Best quant | Size on disk |
|---|---|---|---|---|
| 1 | Qwen3.6-35B-A3B | 35B | UD-Q5_K_XL | 26.6 GB |
| 2 | Qwen2.5-32B-Instruct | 32.764B | Q5_K_M | 23.3 GB |
| 3 | QwQ-32B | 32.764B | Q5_K_M | 23.3 GB |
| 4 | Qwen3-32B | 32.762B | Q5_K_M | 23.2 GB |
| 5 | DeepSeek-R1-Distill-Qwen-32B | 32.76B | Q5_K_M | 23.3 GB |
| 6 | Qwen3-30B-A3B | 30.5B | Q6_K | 25.1 GB |
| 7 | Qwen3.6-27B | 27.78B | Q6_K | 22.5 GB |
| 8 | gemma-3-27b-it | 27.43B | Q6_K | 22.2 GB |
| 9 | gemma-2-27b-it | 27.23B | Q6_K | 22.3 GB |
| 10 | gemma-4-26B-A4B-it | 26.5B | UD-Q5_K_M | 21.1 GB |
| 11 | Mistral-Small-24B-Instruct-2501 | 23.57B | Q8_0 | 25.1 GB |
| 12 | gpt-oss-20b | 21B | MXFP4 | 12.1 GB |
Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.