Ollama · 8GB VRAM
Best Ollama models for 8GB VRAM
13 models we track load cleanly on 8GB — quality quantization, no CPU offload, 8K context. The largest is DeepSeek-R1-Distill-Llama-8B at Q4_K_M. Verdicts are computed on a NVIDIA GeForce RTX 4060, a real card we have sourced specs for — not a generic 8GB abstraction.
Cards with 8GB: NVIDIA GeForce RTX 3070 · NVIDIA GeForce RTX 3070 Ti · NVIDIA GeForce RTX 4060 · NVIDIA GeForce RTX 4060 Ti 8 GB.
| # | Model | Params | Best quant | Size on disk |
|---|---|---|---|---|
| 1 | DeepSeek-R1-Distill-Llama-8B | 8.03B | Q4_K_M | 4.9 GB |
| 2 | Llama-3.1-8B-Instruct | 8.03B | Q4_K_M | 4.9 GB |
| 3 | DeepSeek-R1-Distill-Qwen-7B | 7.62B | Q5_K_M | 5.4 GB |
| 4 | Qwen2.5-7B-Instruct | 7.616B | Q5_K_M | 5.4 GB |
| 5 | Mistral-7B-Instruct-v0.3 | 7.25B | Q4_K_M | 4.4 GB |
| 6 | gemma-3-4b-it | 4.3B | Q8_0 | 4.1 GB |
| 7 | Phi-3.5-mini-instruct | 3.82B | Q5_K_M | 2.8 GB |
| 8 | Llama-3.2-3B-Instruct | 3.213B | Q8_0 | 3.4 GB |
| 9 | Qwen2.5-3B-Instruct | 3.086B | Q8_0 | 3.3 GB |
| 10 | gemma-2-2b-it | 2.61B | Q8_0 | 2.8 GB |
| 11 | Qwen2.5-1.5B-Instruct | 1.544B | Q8_0 | 1.6 GB |
| 12 | Llama-3.2-1B-Instruct | 1.236B | Q8_0 | 1.3 GB |
Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.