TokenAssemble

Ollama · 8GB VRAM

Best Ollama models for 8GB VRAM

13 models we track load cleanly on 8GB — quality quantization, no CPU offload, 8K context. The largest is DeepSeek-R1-Distill-Llama-8B at Q4_K_M. Verdicts are computed on a NVIDIA GeForce RTX 4060, a real card we have sourced specs for — not a generic 8GB abstraction.

Cards with 8GB: NVIDIA GeForce RTX 3070 · NVIDIA GeForce RTX 3070 Ti · NVIDIA GeForce RTX 4060 · NVIDIA GeForce RTX 4060 Ti 8 GB.

#ModelParamsBest quantSize on disk
1DeepSeek-R1-Distill-Llama-8B8.03BQ4_K_M4.9 GB
2Llama-3.1-8B-Instruct8.03BQ4_K_M4.9 GB
3DeepSeek-R1-Distill-Qwen-7B7.62BQ5_K_M5.4 GB
4Qwen2.5-7B-Instruct7.616BQ5_K_M5.4 GB
5Mistral-7B-Instruct-v0.37.25BQ4_K_M4.4 GB
6gemma-3-4b-it4.3BQ8_04.1 GB
7Phi-3.5-mini-instruct3.82BQ5_K_M2.8 GB
8Llama-3.2-3B-Instruct3.213BQ8_03.4 GB
9Qwen2.5-3B-Instruct3.086BQ8_03.3 GB
10gemma-2-2b-it2.61BQ8_02.8 GB
11Qwen2.5-1.5B-Instruct1.544BQ8_01.6 GB
12Llama-3.2-1B-Instruct1.236BQ8_01.3 GB

Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.