TokenAssemble

Ollama · 24GB VRAM

Best Ollama models for 24GB VRAM

28 models we track load cleanly on 24GB — quality quantization, no CPU offload, 8K context. The largest is Qwen3-30B-A3B at Q4_K_M. Verdicts are computed on a NVIDIA GeForce RTX 3090, a real card we have sourced specs for — not a generic 24GB abstraction.

Cards with 24GB: AMD Radeon RX 7900 XTX · NVIDIA GeForce RTX 3090 · NVIDIA GeForce RTX 3090 Ti · NVIDIA GeForce RTX 4090.

#ModelParamsBest quantSize on disk
1Qwen3-30B-A3B30.5BQ4_K_M18.6 GB
2Qwen3.6-27B27.78BQ4_K_M16.8 GB
3gemma-2-27b-it27.23BQ4_K_M16.6 GB
4gemma-4-26B-A4B-it26.5BUD-Q4_K_M16.9 GB
5Mistral-Small-24B-Instruct-250123.57BQ5_K_M16.8 GB
6gpt-oss-20b21BMXFP412.1 GB
7DeepSeek-R1-Distill-Qwen-14B14.77BQ8_015.7 GB
8Qwen2.5-14B-Instruct14.77BQ8_015.7 GB
9Qwen3-14B14.768BQ8_015.7 GB
10Phi-414.66BQ8_015.6 GB
11Mistral-Nemo-Instruct-240712.25BQ8_013.0 GB
12gemma-3-12b-it12.19BQ8_012.5 GB

Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.