Ollama · 12GB VRAM
Best Ollama models for 12GB VRAM
17 models we track load cleanly on 12GB — quality quantization, no CPU offload, 8K context. The largest is Mistral-Nemo-Instruct-2407 at Q4_K_M. Verdicts are computed on a NVIDIA GeForce RTX 3060 12 GB, a real card we have sourced specs for — not a generic 12GB abstraction.
Cards with 12GB: Intel Arc B580 · NVIDIA GeForce RTX 3060 12 GB · NVIDIA GeForce RTX 3080 12 GB · NVIDIA GeForce RTX 3080 Ti · NVIDIA GeForce RTX 4070 · NVIDIA GeForce RTX 4070 SUPER · NVIDIA GeForce RTX 4070 Ti · NVIDIA GeForce RTX 5070.
| # | Model | Params | Best quant | Size on disk |
|---|---|---|---|---|
| 1 | Mistral-Nemo-Instruct-2407 | 12.25B | Q4_K_M | 7.5 GB |
| 2 | Qwen3.5-9B | 9.65B | Q6_K | 7.5 GB |
| 3 | gemma-2-9b-it | 9.24B | Q6_K | 7.6 GB |
| 4 | Qwen3-8B | 8.191B | Q6_K | 6.7 GB |
| 5 | DeepSeek-R1-Distill-Llama-8B | 8.03B | Q8_0 | 8.5 GB |
| 6 | Llama-3.1-8B-Instruct | 8.03B | Q8_0 | 8.5 GB |
| 7 | DeepSeek-R1-Distill-Qwen-7B | 7.62B | Q8_0 | 8.1 GB |
| 8 | Qwen2.5-7B-Instruct | 7.616B | Q8_0 | 8.1 GB |
| 9 | Mistral-7B-Instruct-v0.3 | 7.25B | Q8_0 | 7.7 GB |
| 10 | gemma-3-4b-it | 4.3B | Q8_0 | 4.1 GB |
| 11 | Phi-3.5-mini-instruct | 3.82B | Q8_0 | 4.1 GB |
| 12 | Llama-3.2-3B-Instruct | 3.213B | Q8_0 | 3.4 GB |
Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.