CanMyPCRunAILocal AI compatibility

Ollama Library · qwen35

Can I run nimble?

A 9B decision model from Bespoke Labs for fast, typed classification.

Family maximum: 256K contexttextApache License 2.02 builds

Parameter counts and memory belong to each build below. A family can contain several sizes and tasks; fitting in memory is not a quality ranking.

Calculate memory for a build and context

nimble system requirements

Memory needed at 4K context, counting model weights, KV cache, compute buffers and runtime overhead. Download sizes come from the official Ollama registry.

Download size and memory required for each nimble build
BuildQuantizationDownloadMemory needed
nimble:9b-q4_K_M9B parameters (tag)256K documented maximum contextBuild sourceCatalog checked 2026-10-02Q4_K_M5.2 GB~7.9 GB(est.)
nimble:9b9B parameters (tag)256K documented maximum contextBuild sourceCatalog checked 2026-10-02Q8_08.9 GB~13.1 GB(est.)

Figures marked (est.) use the fallback KV estimate because this build's architecture is missing from our catalog. All totals are estimates, including runtime buffers.

Showing the 2 smallest non-alias builds. Choose any build and change the context. A context above the build's documented maximum is a hypothetical calculation, not a supported setting.

How to run nimble locally

  1. 1.Install Ollama for Windows, macOS or Linux.
  2. 2.Download this exact build in a terminal: ollama pull nimble:9b-q4_K_M
  3. 3.Check the source for the build's task. For text generation, run ollama run nimble:9b-q4_K_M. Base/text builds may need a prompt template; an instruct build is usually the appropriate starting point for a conversation.

Frequently asked questions about nimble

How much VRAM do I need to run nimble?
At 4K context, the smallest published build of nimble (Q4_K_M) needs roughly 7.9 GB of GPU memory once weights, KV cache, compute buffers and runtime overhead are counted. Bigger quantizations and longer context windows need more. It can also run partly or entirely on the CPU using system RAM, more slowly.
Can I run nimble without a dedicated GPU?
A supported CPU runtime can use system RAM. The smallest listed build needs an estimated 7.9 GB at 4K context, plus headroom for your operating system. Check the build's runtime requirements. We have not measured its speed on your CPU.
How big is the nimble download?
The smallest published build is 5.2 GB. There are 2 builds in total, the largest being 8.9 GB. Leave some extra free disk space beyond the download itself.
How do I run nimble locally?
Install Ollama, then download a specific build with "ollama pull nimble:9b-q4_K_M". For a build that supports text generation, use ollama run with the same tag. Check the model source for its task and prompt format.

Will nimble run on your PC?

Scan your hardware once and get a verdict for this model and every build — free, no account.

Check my PC

Model data from the official Ollama registry · verified 02/10/2026