What AI models can you run with 32GB of VRAM?
With 32GB of VRAM, 194 of 228 models from the Ollama library run entirely on the GPU at 4K context. 10 more run by offloading part of the model to system RAM, which works but is slower.
194
run well
10
with trade-offs
228
models tested
Best AI models for 32GB of VRAM
Heaviest first — each of these fits entirely in 32GB at 4K context, counting weights, KV cache, compute buffers and runtime overhead.
| Model | Build | Memory needed | Verdict |
|---|---|---|---|
| laguna-xs-2.1 | latest | ~27.4 GB | Good fit |
| codebooga | 34B | ~27.3 GB | Good fit |
| phind-codellama | 34B | ~27.3 GB | Good fit |
| qwq | 32B | ~26.8 GB | Good fit |
| command-r | 35B | ~26.8 GB | Good fit |
| olmo-3.1 | 32B | ~26.3 GB | Good fit |
| glm-4.7-flash | latest | ~25.7 GB | Good fit |
| north-mini-code-1.0 | latest | ~25.2 GB | Good fit |
| qwen3-coder | 30B | ~25.1 GB | Good fit |
| qwen3.6 | 27B | ~23.6 GB | Good fit |
| qwen3.8 | 27B | ~22.8 GB | Excellent fit |
| muse-glimmer | 30B | ~22.7 GB | Excellent fit |
| mistral-small3.1 | 24B | ~21.0 GB | Excellent fit |
| devstral-small-2 | 24B | ~20.6 GB | Excellent fit |
| mistral-small3.2 | 24B | ~20.6 GB | Excellent fit |
| lfm2 | 24B | ~19.6 GB | Excellent fit |
| devstral | 24B | ~19.5 GB | Excellent fit |
| magistral | 24B | ~19.5 GB | Excellent fit |
| gpt-oss | 20B | ~18.8 GB | Excellent fit |
| gpt-oss-safeguard | 20B | ~18.8 GB | Excellent fit |
| mistral-small | 22B | ~18.2 GB | Excellent fit |
| codestral | 22B | ~18.2 GB | Excellent fit |
| solar-pro | 22B | ~18.1 GB | Excellent fit |
| phi4-reasoning | 14B | ~15.2 GB | Excellent fit |
| deepseek-coder-v2 | 16B | ~14.2 GB | Excellent fit |
Graphics cards with 32GB of VRAM
How this was calculated
Every model is evaluated at 4K context against 32GB of VRAM, minus a reserve for the display and operating system, and assuming 32GB of system RAM with working GPU drivers. Your own machine may differ, which is what the PC scan measures directly.
32GB VRAM and local AI — frequently asked questions
- Is 32GB of VRAM enough to run AI models locally?
- Yes — 194 of the 228 models tested run entirely on the GPU with 32GB at 4K context, and 10 more run with part of the model offloaded to system RAM. The heaviest that fits is laguna-xs-2.1:latest, needing about 27.4 GB.
- What is the best AI model for 32GB VRAM?
- For maximum capability, laguna-xs-2.1 (laguna-xs-2.1:latest) is the heaviest build that still fits entirely in 32GB. Smaller models leave headroom for a longer context window, which often matters more than raw parameter count.
- Which graphics cards have 32GB of VRAM?
- Cards with 32GB include GeForce RTX 5090.