TokenAssemble

Ollama · 12GB VRAM

Best Ollama models for 12GB VRAM

17 models we track load cleanly on 12GB — quality quantization, no CPU offload, 8K context. The largest is Mistral-Nemo-Instruct-2407 at Q4_K_M. Verdicts are computed on a NVIDIA GeForce RTX 3060 12 GB, a real card we have sourced specs for — not a generic 12GB abstraction.

Cards with 12GB: Intel Arc B580 · NVIDIA GeForce RTX 3060 12 GB · NVIDIA GeForce RTX 3080 12 GB · NVIDIA GeForce RTX 3080 Ti · NVIDIA GeForce RTX 4070 · NVIDIA GeForce RTX 4070 SUPER · NVIDIA GeForce RTX 4070 Ti · NVIDIA GeForce RTX 5070.

#ModelParamsBest quantSize on disk
1Mistral-Nemo-Instruct-240712.25BQ4_K_M7.5 GB
2Qwen3.5-9B9.65BQ6_K7.5 GB
3gemma-2-9b-it9.24BQ6_K7.6 GB
4Qwen3-8B8.191BQ6_K6.7 GB
5DeepSeek-R1-Distill-Llama-8B8.03BQ8_08.5 GB
6Llama-3.1-8B-Instruct8.03BQ8_08.5 GB
7DeepSeek-R1-Distill-Qwen-7B7.62BQ8_08.1 GB
8Qwen2.5-7B-Instruct7.616BQ8_08.1 GB
9Mistral-7B-Instruct-v0.37.25BQ8_07.7 GB
10gemma-3-4b-it4.3BQ8_04.1 GB
11Phi-3.5-mini-instruct3.82BQ8_04.1 GB
12Llama-3.2-3B-Instruct3.213BQ8_03.4 GB

Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.