Ollama · 48GB VRAM
Best Ollama models for 48GB VRAM
34 models we track load cleanly on 48GB — quality quantization, no CPU offload, 8K context. The largest is Qwen3.6-35B-A3B at Q8_0. Verdicts are computed on a NVIDIA RTX A6000, a real card we have sourced specs for — not a generic 48GB abstraction.
Cards with 48GB: AMD Radeon PRO W7900 · NVIDIA RTX 6000 Ada Generation · NVIDIA RTX A6000.
| # | Model | Params | Best quant | Size on disk |
|---|---|---|---|---|
| 1 | Qwen3.6-35B-A3B | 35B | Q8_0 | 36.9 GB |
| 2 | Qwen2.5-32B-Instruct | 32.764B | Q8_0 | 34.8 GB |
| 3 | QwQ-32B | 32.764B | Q8_0 | 34.8 GB |
| 4 | Qwen3-32B | 32.762B | Q8_0 | 34.8 GB |
| 5 | DeepSeek-R1-Distill-Qwen-32B | 32.76B | Q8_0 | 34.8 GB |
| 6 | Qwen3-30B-A3B | 30.5B | Q8_0 | 32.5 GB |
| 7 | Qwen3.6-27B | 27.78B | Q8_0 | 28.6 GB |
| 8 | gemma-3-27b-it | 27.43B | Q8_0 | 28.7 GB |
| 9 | gemma-2-27b-it | 27.23B | Q8_0 | 28.9 GB |
| 10 | gemma-4-26B-A4B-it | 26.5B | Q8_0 | 26.9 GB |
| 11 | Mistral-Small-24B-Instruct-2501 | 23.57B | Q8_0 | 25.1 GB |
| 12 | gpt-oss-20b | 21B | MXFP4 | 12.1 GB |
Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.