LLM VRAM calculator
A model download is only part of its memory cost. Choose a concrete build and conversation length to estimate weights, KV cache, runtime overhead and compute buffers. No hardware profile or download is required.
Estimated memory: 7.9 GB
nimble:9b-q4_K_M · Q4_K_M · 4,096 tokens · ollama
- Model weights
- 5.2 GB
- Context / KV cache
- 1.8 GB
- Runtime overhead
- 0.4 GB
- Compute buffers
- 0.4 GB
- Total estimate
- 7.9 GB
Low confidence: architecture data is missing, so the KV cache uses a fallback estimate. These are working-memory estimates, not a VRAM-capacity guarantee. Your display and operating system need headroom. On Apple Silicon, CPU and GPU share the same memory pool.
Download: 5.2 GB. Build size from tag: 9B parameters. GB here means 1,024³ bytes.
Build manifest · Catalog checked 2026-10-02.
Download this exact build
ollama pull nimble:9b-q4_K_MFor a generation-capable build, run ollama run nimble:9b-q4_K_M, then enter /set parameter num_ctx 4096 before your prompt. Check the source for base versus instruct behavior.
How to use the number
Compare quantizations of the same build at the same context. A lower-bit build often uses less memory but can change output quality. Longer context grows the KV cache for autoregressive models. Encoder-only models can have different cache behavior. This calculator describes inference, not training or fine-tuning.
The engine estimates one loaded model. Parallel requests, vision inputs, other applications and runtime settings can raise actual use. A fit does not predict tokens per second.