TokenAssemble

Ollama · 48GB VRAM

Best Ollama models for 48GB VRAM

34 models we track load cleanly on 48GB — quality quantization, no CPU offload, 8K context. The largest is Qwen3.6-35B-A3B at Q8_0. Verdicts are computed on a NVIDIA RTX A6000, a real card we have sourced specs for — not a generic 48GB abstraction.

Cards with 48GB: AMD Radeon PRO W7900 · NVIDIA RTX 6000 Ada Generation · NVIDIA RTX A6000.

#ModelParamsBest quantSize on disk
1Qwen3.6-35B-A3B35BQ8_036.9 GB
2Qwen2.5-32B-Instruct32.764BQ8_034.8 GB
3QwQ-32B32.764BQ8_034.8 GB
4Qwen3-32B32.762BQ8_034.8 GB
5DeepSeek-R1-Distill-Qwen-32B32.76BQ8_034.8 GB
6Qwen3-30B-A3B30.5BQ8_032.5 GB
7Qwen3.6-27B27.78BQ8_028.6 GB
8gemma-3-27b-it27.43BQ8_028.7 GB
9gemma-2-27b-it27.23BQ8_028.9 GB
10gemma-4-26B-A4B-it26.5BQ8_026.9 GB
11Mistral-Small-24B-Instruct-250123.57BQ8_025.1 GB
12gpt-oss-20b21BMXFP412.1 GB

Model names in Ollama's registry may differ from the names above — we link each model to its source repository rather than guess a pull tag. See the Ollama guide for setup and GPU requirements, or check your exact card.