What AI models can you run with 4GB of VRAM?
With 4GB of VRAM, 68 of 228 models from the Ollama library run entirely on the GPU at 4K context. 126 more run by offloading part of the model to system RAM, which works but is slower.
68
run well
126
with trade-offs
228
models tested
Best AI models for 4GB of VRAM
Heaviest first — each of these fits entirely in 4GB at 4K context, counting weights, KV cache, compute buffers and runtime overhead.
| Model | Build | Memory needed | Verdict |
|---|---|---|---|
| granite4.2 | 3B | ~3.4 GB | Good fit |
| cogito | 3B | ~3.4 GB | Good fit |
| glm-ocr | latest | ~3.3 GB | Good fit |
| granite-code | 3B | ~3.2 GB | Good fit |
| smallthinker | 3B | ~3.2 GB | Good fit |
| granite4 | 3B | ~3.2 GB | Good fit |
| granite4.1 | 3B | ~3.2 GB | Good fit |
| hermes3 | 3B | ~3.1 GB | Good fit |
| qwen3.5 | 2B | ~3.0 GB | Good fit |
| qwen3-vl | 2B | ~2.9 GB | Good fit |
| starcoder2 | 3B | ~2.9 GB | Good fit |
| dolphin-phi | 2.7B | ~2.8 GB | Good fit |
| phi | 2.7B | ~2.8 GB | Good fit |
| stablelm-zephyr | 3B | ~2.7 GB | Excellent fit |
| stable-code | 3B | ~2.7 GB | Excellent fit |
| shieldgemma | 2B | ~2.7 GB | Excellent fit |
| gemma2 | 2B | ~2.7 GB | Excellent fit |
| exaone-deep | 2.4B | ~2.6 GB | Excellent fit |
| exaone3.5 | 2.4B | ~2.6 GB | Excellent fit |
| gemma | 2B | ~2.6 GB | Excellent fit |
| codegemma | 2B | ~2.6 GB | Excellent fit |
| granite3-dense | 2B | ~2.5 GB | Excellent fit |
| granite3.1-dense | 2B | ~2.5 GB | Excellent fit |
| granite3.3 | 2B | ~2.5 GB | Excellent fit |
| granite3.2 | 2B | ~2.5 GB | Excellent fit |
How this was calculated
Every model is evaluated at 4K context against 4GB of VRAM, minus a reserve for the display and operating system, and assuming 32GB of system RAM with working GPU drivers. Your own machine may differ, which is what the PC scan measures directly.
4GB VRAM and local AI — frequently asked questions
- Is 4GB of VRAM enough to run AI models locally?
- Yes — 68 of the 228 models tested run entirely on the GPU with 4GB at 4K context, and 126 more run with part of the model offloaded to system RAM. The heaviest that fits is granite4.2:3b, needing about 3.4 GB.
- What is the best AI model for 4GB VRAM?
- For maximum capability, granite4.2 (granite4.2:3b) is the heaviest build that still fits entirely in 4GB. Smaller models leave headroom for a longer context window, which often matters more than raw parameter count.
- Which graphics cards have 4GB of VRAM?
- Several cards ship with 4GB of VRAM across NVIDIA and AMD ranges.