Ollama · 24GB VRAM
Best Ollama models for 24GB VRAM
28 models we track load cleanly on 24GB — quality quantization, no CPU offload, 8K context. The largest is Qwen3-30B-A3B at Q4_K_M. Verdicts are computed on a NVIDIA GeForce RTX 3090, a real card we have sourced specs for — not a generic 24GB abstraction.
Cards with 24GB: AMD Radeon RX 7900 XTX · NVIDIA GeForce RTX 3090 · NVIDIA GeForce RTX 3090 Ti · NVIDIA GeForce RTX 4090.
| # | Model | Params | Best quant | Size on disk |
|---|---|---|---|---|
| 1 | Qwen3-30B-A3B | 30.5B | Q4_K_M | 18.6 GB |
| 2 | Qwen3.6-27B | 27.78B | Q4_K_M | 16.8 GB |
| 3 | gemma-2-27b-it | 27.23B | Q4_K_M | 16.6 GB |
| 4 | gemma-4-26B-A4B-it | 26.5B | UD-Q4_K_M | 16.9 GB |
| 5 | Mistral-Small-24B-Instruct-2501 | 23.57B | Q5_K_M | 16.8 GB |
| 6 | gpt-oss-20b | 21B | MXFP4 | 12.1 GB |
| 7 | DeepSeek-R1-Distill-Qwen-14B | 14.77B | Q8_0 | 15.7 GB |
| 8 | Qwen2.5-14B-Instruct | 14.77B | Q8_0 | 15.7 GB |
| 9 | Qwen3-14B | 14.768B | Q8_0 | 15.7 GB |
| 10 | Phi-4 | 14.66B | Q8_0 | 15.6 GB |
| 11 | Mistral-Nemo-Instruct-2407 | 12.25B | Q8_0 | 13.0 GB |
| 12 | gemma-3-12b-it | 12.19B | Q8_0 | 12.5 GB |
Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.