What AI models can you run with 8GB of VRAM?
With 8GB of VRAM, 149 of 228 models from the Ollama library run entirely on the GPU at 4K context. 50 more run by offloading part of the model to system RAM, which works but is slower.
149
run well
50
with trade-offs
228
models tested
Best AI models for 8GB of VRAM
Heaviest first — each of these fits entirely in 8GB at 4K context, counting weights, KV cache, compute buffers and runtime overhead.
| Model | Build | Memory needed | Verdict |
|---|---|---|---|
| lfm2.5 | 8B | ~7.3 GB | Good fit |
| granite4.1-guardian | 8B | ~7.2 GB | Good fit |
| rnj-1 | 8B | ~7.2 GB | Good fit |
| command-r7b-arabic | 7B | ~7.1 GB | Good fit |
| command-r7b | 7B | ~7.1 GB | Good fit |
| aya-expanse | 8B | ~7.1 GB | Good fit |
| llama-pro | 8B | ~7.1 GB | Good fit |
| minicpm-v4.5 | 8B | ~7.1 GB | Good fit |
| llava-llama3 | 8B | ~6.9 GB | Good fit |
| tulu3 | 8B | ~6.9 GB | Good fit |
| llama3-groq-tool-use | 8B | ~6.9 GB | Good fit |
| dolphin3 | 8B | ~6.9 GB | Good fit |
| dolphin-llama3 | 8B | ~6.9 GB | Good fit |
| llama3.1 | 8B | ~6.9 GB | Good fit |
| llama3-gradient | 8B | ~6.9 GB | Good fit |
| llama3 | 8B | ~6.9 GB | Good fit |
| llama3-chatqa | 8B | ~6.9 GB | Good fit |
| aya | 8B | ~6.8 GB | Good fit |
| codeqwen | 7B | ~6.7 GB | Good fit |
| bespoke-minicheck | 7B | ~6.7 GB | Good fit |
| openthinker | 7B | ~6.6 GB | Good fit |
| marco-o1 | 7B | ~6.6 GB | Good fit |
| minicpm-v | 8B | ~6.6 GB | Good fit |
| dolphincoder | 7B | ~6.5 GB | Good fit |
| olmo2 | 7B | ~6.3 GB | Good fit |
Graphics cards with 8GB of VRAM
How this was calculated
Every model is evaluated at 4K context against 8GB of VRAM, minus a reserve for the display and operating system, and assuming 32GB of system RAM with working GPU drivers. Your own machine may differ, which is what the PC scan measures directly.
8GB VRAM and local AI — frequently asked questions
- Is 8GB of VRAM enough to run AI models locally?
- Yes — 149 of the 228 models tested run entirely on the GPU with 8GB at 4K context, and 50 more run with part of the model offloaded to system RAM. The heaviest that fits is lfm2.5:8b, needing about 7.3 GB.
- What is the best AI model for 8GB VRAM?
- For maximum capability, lfm2.5 (lfm2.5:8b) is the heaviest build that still fits entirely in 8GB. Smaller models leave headroom for a longer context window, which often matters more than raw parameter count.
- Which graphics cards have 8GB of VRAM?
- Cards with 8GB include GeForce RTX 5060, GeForce RTX 4060 Ti 8GB, GeForce RTX 4060, GeForce RTX 3070, GeForce RTX 3060 Ti, GeForce RTX 3050 8GB, Radeon RX 7600.