TokenAssemble

Ollama · 32GB VRAM

Best Ollama models for 32GB VRAM

34 models we track load cleanly on 32GB — quality quantization, no CPU offload, 8K context. The largest is Qwen3.6-35B-A3B at UD-Q5_K_XL. Verdicts are computed on a NVIDIA GeForce RTX 5090, a real card we have sourced specs for — not a generic 32GB abstraction.

Cards with 32GB: NVIDIA GeForce RTX 5090 · NVIDIA RTX 5000 Ada Generation.

#ModelParamsBest quantSize on disk
1Qwen3.6-35B-A3B35BUD-Q5_K_XL26.6 GB
2Qwen2.5-32B-Instruct32.764BQ5_K_M23.3 GB
3QwQ-32B32.764BQ5_K_M23.3 GB
4Qwen3-32B32.762BQ5_K_M23.2 GB
5DeepSeek-R1-Distill-Qwen-32B32.76BQ5_K_M23.3 GB
6Qwen3-30B-A3B30.5BQ6_K25.1 GB
7Qwen3.6-27B27.78BQ6_K22.5 GB
8gemma-3-27b-it27.43BQ6_K22.2 GB
9gemma-2-27b-it27.23BQ6_K22.3 GB
10gemma-4-26B-A4B-it26.5BUD-Q5_K_M21.1 GB
11Mistral-Small-24B-Instruct-250123.57BQ8_025.1 GB
12gpt-oss-20b21BMXFP412.1 GB

Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.