TokenAssemble

AMD GPU

AMD Radeon RX 7900 XT

The AMD Radeon RX 7900 XT has 20GB of VRAM and 800 GB/s of memory bandwidth. Its sweet spot tops out around Mistral-Small-24B-Instruct-2501 — the largest model that loads cleanly at a quality quant with llama.cpp at 8K context.

AMD verdicts are beta: fit is physics and holds, but ROCm/Vulkan speed tiers aren't calibrated yet — treat tiers as provisional until real measurements land.
VRAM
20 GB
Memory bandwidth
800 GB/s
TDP
315 W
MSRP
$899
Backends
ROCm · Vulkan

Also good for gaming — 4K-class (estimated from FP32 compute, not a game benchmark)

Specs as of 2026-07-08 · source

AMD Radeon RX 7900 XT for local AI — common questions

How much VRAM does the AMD Radeon RX 7900 XT have?
The AMD Radeon RX 7900 XT has 20GB of VRAM with 800 GB/s of memory bandwidth. VRAM sets which models fit; bandwidth sets how fast they feel once they do.
Is the AMD Radeon RX 7900 XT good for running AI models locally?
Yes, within limits. 20GB comfortably runs small and mid-size models at a quality quant, but the largest open models will not fit. The largest model it loads cleanly is Mistral-Small-24B-Instruct-2501, at a quality quant with llama.cpp at 8K context.
What runtimes work with the AMD Radeon RX 7900 XT?
It runs on ROCm · Vulkan, which covers Ollama, llama.cpp and LM Studio. Runtime choice affects speed and features, not whether a model fits.

Affiliate link — we may earn a commission. Verdicts are computed before any link is attached.

What an AMD Radeon RX 7900 XT can run

Every model we track, graded at its best clean-fit quant. Click any verdict for the full breakdown.

ModelBest-quant verdict · llama.cpp · 8K context
Qwen3-8BFits · A · Q8_0interactivebeta
gemma-4-26B-A4B-itTight · C · MXFP4_MOEinteractivebeta
Qwen2.5-7B-InstructFits · A · Q8_0interactivebeta
Qwen2.5-1.5B-InstructFits · A · Q8_0interactivebeta
Llama-3.2-1B-InstructFits · A · Q8_0interactivebeta
Llama-3.1-8B-InstructFits · A · Q8_0interactivebeta
DeepSeek-R1Won't fitbeta
Qwen3.5-9BFits · A · Q8_0interactivebeta
gpt-oss-20bFits · A · MXFP4interactivebeta
Qwen3.6-35B-A3BTight · C · MXFP4_MOEinteractive · offloadbeta
Qwen2.5-3B-InstructFits · A · Q8_0interactivebeta
Qwen3-32BTight · C · Q4_K_Musable · offloadbeta
Qwen3.6-27BTight · C · Q4_K_Minteractivebeta
Qwen2.5-0.5B-InstructFits · A · Q8_0interactivebeta
Mistral-7B-Instruct-v0.3Fits · A · Q8_0interactivebeta
gpt-oss-120bTight · C · MXFP4interactive · offloadbeta
Qwen3-14BFits · B · Q6_Kinteractivebeta
Qwen3-30B-A3BTight · C · Q4_K_Minteractive · offloadbeta
Qwen2.5-32B-InstructTight · C · Q4_K_Musable · offloadbeta
DeepSeek-V4-FlashWon't fitbeta
Qwen2.5-14B-InstructFits · B · Q6_Kinteractivebeta
Llama-3.2-3B-InstructFits · A · Q8_0interactivebeta
gemma-3-4b-itFits · A · Q8_0interactivebeta
gemma-3-12b-itFits · B · Q8_0interactivebeta
Llama-3.1-70B-InstructTight · C · Q4_K_Mpainful · offloadbeta
Phi-3.5-mini-instructFits · A · Q8_0interactivebeta
Phi-4Fits · B · Q6_Kinteractivebeta
DeepSeek-R1-Distill-Qwen-32BTight · C · Q4_K_Musable · offloadbeta
gemma-3-27b-itTight · C · Q4_K_Musable · offloadbeta
Llama-3.3-70B-InstructTight · C · Q4_K_Mpainful · offloadbeta
Qwen2.5-72B-InstructTight · C · Q4_K_Mpainful · offloadbeta
gemma-2-2b-itFits · A · Q8_0interactivebeta
DeepSeek-R1-Distill-Qwen-14BFits · B · Q6_Kinteractivebeta
DeepSeek-R1-Distill-Llama-8BFits · A · Q8_0interactivebeta
DeepSeek-R1-Distill-Qwen-7BFits · A · Q8_0interactivebeta
gemma-2-9b-itFits · A · Q8_0interactivebeta
Mistral-Nemo-Instruct-2407Fits · B · Q8_0interactivebeta
QwQ-32BTight · C · Q4_K_Musable · offloadbeta
Mistral-Small-24B-Instruct-2501Fits · B · Q4_K_Minteractivebeta
gemma-2-27b-itTight · C · Q4_K_Minteractivebeta

Have a specific model in mind? Check it against the AMD Radeon RX 7900 XT at your own quant and context length.

Check with the AMD Radeon RX 7900 XT preselected →

Verdicts computed from published specs and measured GGUF sizes — how the math works.