Downloading a large local AI model only to discover that your GPU cannot actually run it is an easy mistake to make.
A model may be advertised as an “8B” model, but parameter count alone doesn't tell you how much GPU memory inference will require.
The memory used during inference depends on several things:
- Model weights
- Quantization format
- Context length
- KV cache
- Batch or concurrent requests
- Inference framework
- GPU memory already being used by other applications
That is why an 8B model can fit comfortably in one configuration and still run out of memory in another.
The useful question isn't:
“How large is the model?”
It's:
“How much memory will this particular model configuration need at my target context length?”
This guide explains how to estimate that before downloading a model.
The Basic Inference Memory Equation
A simplified model of GPU memory usage is:
Total GPU Memory ≈ Model Weights + KV Cache + Runtime/Activation Memory + Other GPU Usage
This is an estimate rather than an exact universal formula.
Different inference engines allocate memory differently, and features such as batching, parallel requests, Flash Attention, quantized KV caches, and CPU offloading can change the final number.
So the goal of the calculation is not to predict the exact number shown by nvidia-smi.
The goal is to determine whether you have enough headroom to run the configuration you actually want.
1. Start With Model Weight Memory
The first component is the model's weights.
A simple lower-bound estimate is:
Weight Memory ≈ Parameter Count × Bits per Parameter ÷ 8
For example, an 8-billion-parameter model stored at approximately 4 bits per parameter would require:
8B × 4 ÷ 8 ≈ 4 GB
But that does not mean every 4-bit 8B model will occupy exactly 4 GB.
Quantized model files contain additional information, and different quantization schemes have different overheads.
The actual model file size is therefore a better starting point than multiplying the parameter count by a single fixed multiplier.
Approximate Weight-Only Estimates
A useful first-pass calculation looks like this:
| Model | Approximate raw weight calculation |
|---|---|
| 7B at FP16 | ~14 GB |
| 8B at FP16 | ~16 GB |
| 14B at FP16 | ~28 GB |
| 32B at FP16 | ~64 GB |
| 70B at FP16 | ~140 GB |
For lower-bit quantization, the theoretical weight requirement falls substantially.
For example:
| Model | 8-bit | 4-bit theoretical minimum |
|---|---|---|
| 8B | ~8 GB | ~4 GB |
| 14B | ~14 GB | ~7 GB |
| 32B | ~32 GB | ~16 GB |
| 70B | ~70 GB | ~35 GB |
These are planning estimates, not exact download sizes or guaranteed VRAM requirements.
For GGUF models in particular, check the actual file size and quantization type before downloading.
2. The Part Many People Forget: KV Cache
Model weights are relatively static.
The KV cache is different.
During generation, the model stores attention information from previously processed tokens so it doesn't have to recompute the same information repeatedly.
As the context grows, the KV cache grows too.
This is why a model that works at 4K context can run out of memory at 32K or 128K context.
How to Calculate KV Cache Memory
For a standard transformer using grouped-query attention, a useful simplified formula is:
KV Cache per Token = 2 × Layers × KV Heads × Head Dimension × Bytes per Element
The 2 represents the Key and Value tensors.
For example, Llama 3.1 8B has:
- 32 transformer layers
- 8 key/value heads
- 128-dimensional attention heads
These architecture values are documented in the model configuration. (Hugging Face)
Assuming FP16 KV storage, where each value uses 2 bytes:
2 × 32 × 8 × 128 × 2
= 131,072 bytes per token
That's approximately:
128 KiB per token
This gives us a useful illustration of how context length affects memory.
Llama 3.1 8B KV Cache Example
Using the simplified FP16 calculation above:
| Context | Approximate KV Cache |
|---|---|
| 2K tokens | ~256 MiB |
| 8K tokens | ~1 GiB |
| 32K tokens | ~4 GiB |
| 128K tokens | ~16 GiB |
The important point is the growth.
At 8K context, the KV cache in this simplified example is around 1 GiB.
At 32K, it is around 4 GiB.
At 128K, it reaches around 16 GiB.
That means a model can fit comfortably when running a short context and still run out of memory when you start feeding it long documents.
Actual memory usage can differ because inference engines may use different KV layouts, cache precisions, attention implementations, batching strategies, and other optimizations.
Why Context Length Matters So Much
Consider an 8B model whose quantized weights occupy roughly 5 GB.
You might have a 12 GB GPU and conclude:
12 GB - 5 GB = 7 GB remaining.
That looks comfortable.
But the remaining memory also needs to accommodate the KV cache and runtime allocations.
At a sufficiently long context, the KV cache can consume several additional gigabytes.
And if another application is already using GPU memory, your available headroom becomes even smaller.
This is why model size alone isn't enough to determine whether an LLM will fit.
3. Runtime Memory Also Matters
The remaining GPU memory isn't simply “free space.”
Your inference framework needs memory for additional operations and data structures.
Depending on the framework and workload, this can include:
- Temporary tensors
- Attention operations
- CUDA allocations
- Model buffers
- CUDA graphs or related runtime structures
- Batch processing
- Multiple concurrent sequences
- Framework-specific memory management
The exact amount varies considerably.
For example, vLLM exposes GPU-memory controls such as gpu_memory_utilization and has its own KV-cache memory management. Its current documentation lists a default gpu_memory_utilization value of 0.92. (vLLM)
This is one reason a universal “add exactly 1 GB” rule should not be treated as a precise engineering calculation.
4. Your Desktop May Already Be Using VRAM
If you're running a local model on a desktop GPU, the operating system and other applications may already be using some GPU memory.
Your browser, games, video applications, desktop compositor, IDE, and other GPU-accelerated applications can all consume resources.
Windows also manages dedicated and system GPU memory through its graphics memory manager rather than treating VRAM as a completely isolated pool available to one application. (Microsoft Learn)
So if your GPU has:
12 GB VRAM
you shouldn't assume that all 12 GB will be available to the inference engine.
The amount available depends on what else is running.
A Better Practical Formula
Instead of claiming an exact universal formula, use this as a planning model:
Required GPU Memory ≈ Weight Memory + KV Cache + Runtime Overhead + Safety Headroom
For example:
Model weights ≈ 5 GB
KV cache ≈ 2 GB
Runtime overhead ≈ 1 GB+
Safety headroom ≈ 1–2 GB
--------------------------------
Planning requirement ≈ 9–10 GB+This doesn't guarantee that the model will run.
It tells you that a 12 GB GPU is in a more comfortable position than an 8 GB GPU for that particular configuration.
A Worked Example: 8B Model on a 12 GB GPU
Suppose you have:
GPU: 12 GB VRAM
Model: 8B
Quantization: approximately 4-bit
Target context: 8K
Start with the model weights.
Assume the downloaded quantized model occupies roughly:
~5 GB
Then estimate the KV cache.
For the Llama 3.1 8B example at FP16 KV precision:
8K context ≈ 1 GiB of KV cache
Now you still need memory for runtime allocations and whatever your GPU is already using.
A rough planning model might therefore look like:
5 GB Model
~1 GiB KV cache
+ Runtime allocations
+ Existing GPU usage
+ Safety margin
---------------------------This could fit on a 12 GB GPU, but whether it actually does depends on the inference engine, operating environment, model file, context configuration, and other GPU workloads.
That's much more useful than simply saying:
“8B models need 8 GB.”
They don't.
Quantization Is Only Part of the Equation
Quantization reduces the memory required for model weights, but it doesn't eliminate the KV cache.
For example:
Higher precision model
↓
Larger weights
↓
More VRAM
Lower precision model
↓
Smaller weights
↓
More VRAM available for contextThis is one reason quantized models are so useful for local inference.
However, if your context becomes very large, the KV cache can become a significant part of the total memory requirement even when the model weights are heavily quantized.
KV Cache Quantization
Some inference engines can also reduce the memory used by the KV cache.
Current llama.cpp supports configurable KV-cache types through options such as:
-ctk q8_0
-ctv q8_0for the Key and Value cache respectively. Its current server documentation also supports several other KV-cache types. (GitHub)
For example:
llama-server \
-m model.gguf \
-ctk q8_0 \
-ctv q8_0 \
-c 16384This can reduce KV-cache memory compared with FP16 storage.
The exact memory savings and quality impact depend on the model and workload, so treat KV quantization as an optimization to benchmark rather than assuming that it is completely lossless.
vLLM Also Supports KV Cache Quantization
vLLM currently supports several KV-cache data types, including FP8 variants. (vLLM)
For example:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--kv-cache-dtype fp8vLLM documents FP8 KV-cache quantization as a way to reduce KV-cache memory and support larger token capacity. (vLLM)
Again, compatibility depends on your hardware, model, backend, and vLLM version.
5. Set a Realistic Context Window
One of the easiest ways to control memory usage is to avoid requesting a context window larger than your application actually needs.
With llama-server, the -c / --ctx-size option controls the prompt context size. The current documentation also notes that the default can come from the model configuration when set to zero. (GitHub)
For example:
llama-server \
-m model.gguf \
-c 8192If your application only needs approximately 8K tokens, there's little reason to configure it for 128K simply because the model supports that maximum.
Remember:
Maximum supported context ≠ context you should always allocate.
6. CPU Offloading Can Help
If the complete model doesn't fit into GPU memory, some inference engines can place part of the workload in system RAM.
With llama.cpp, for example, GPU layer offloading can be controlled with options such as -ngl. (GitHub)
A simplified example is:
llama-server \
-m model.gguf \
-ngl 24This doesn't magically make the model as fast as a fully GPU-resident model.
Instead, it trades performance for the ability to run a model that would otherwise exceed the available GPU memory.
The exact number of layers that fit depends on the model, quantization, GPU, context configuration, and other allocations.
What Different VRAM Sizes Actually Mean
Instead of assigning one fixed model size to each GPU, think in terms of workload ranges.
8 GB VRAM
A practical target is generally smaller quantized models.
You can experiment with:
- 3B–8B models
- 4-bit quantization
- Moderate context lengths
- CPU offloading when necessary
An 8B 4-bit model may be possible, but you need to account for KV cache and runtime memory.
Don't assume that the entire 8 GB is available to the model.
12 GB VRAM
This gives you more flexibility for local development.
Typical possibilities include:
- 7B–8B models at higher quantization levels
- 12B–14B models at lower quantization levels
- Moderate context windows
- Some CPU/GPU hybrid configurations
Whether a particular 14B model fits depends on its actual quantized file size and the context you configure.
16 GB VRAM
16 GB opens up more options for local experimentation.
You can consider:
- 7B–14B models at relatively high precision
- Some 30B-class models with aggressive quantization
- Longer context configurations for smaller models
- More room for runtime overhead
But again, the model's actual architecture and quantization matter more than the parameter count alone.
24 GB VRAM
24 GB is a useful capacity for serious local experimentation.
Depending on the model and quantization, you can explore:
- 14B-class models with relatively high precision
- 30B–35B-class models using lower-bit quantization
- Longer context configurations
- Larger models using partial CPU offloading
A 70B model can still be far beyond the comfortable capacity of a single 24 GB GPU, especially at higher precision.
Apple Silicon Is Different
Apple Silicon systems use unified memory rather than a conventional discrete GPU VRAM pool.
That means the CPU and GPU share system memory.
A Mac with 64 GB of unified memory therefore isn't directly comparable to a desktop GPU with 64 GB of dedicated VRAM.
The amount available to a model depends on the operating system, application, model runtime, and other processes.
For Apple Silicon users, the practical question is therefore:
How much unified memory remains available for the inference workload?
rather than simply:
How much VRAM does my GPU have?
A Simple Pre-Download Calculation
Before downloading a model, collect these numbers:
1. Model size
Check the actual quantized model file.
For example:
Q4 model
≈ 5 GB2. Context length
Decide what you actually need:
4K
8K
16K
32KDon't automatically select the model's maximum context.
3. KV cache
Estimate KV memory using the model architecture:
KV Bytes/Token = 2 × Layers × KV Heads × Head Dimension × Bytes/Element
Then multiply by the number of tokens you want to support.
4. Runtime headroom
Leave space for:
- Inference engine allocations
- CUDA/runtime memory
- Other GPU applications
- Batch processing
- Multiple concurrent requests
5. Safety margin
Don't plan for a configuration that uses essentially 100% of your GPU memory.
A little headroom makes troubleshooting considerably easier.
A Useful VRAM Planning Table
| GPU Memory | Practical Starting Point |
|---|---|
| 8 GB | Small-to-medium quantized models |
| 12 GB | 7B–14B quantized models |
| 16 GB | Larger 14B-class models and some aggressively quantized larger models |
| 24 GB | 14B–35B-class models depending on quantization and context |
| 32 GB+ | More room for larger models, longer contexts, and higher precision |
These are planning ranges, not hard limits.
A 12 GB GPU can run a model that a generic chart says requires more memory if some layers are offloaded to system RAM.
Likewise, a model that theoretically fits inside 12 GB can still fail if the context window or runtime allocations push total usage beyond the available memory.
Three Common VRAM Mistakes
Mistake 1: Looking Only at the Download Size
A 5 GB model file does not mean the application will use exactly 5 GB of VRAM.
The runtime needs additional memory.
Mistake 2: Ignoring Context Length
An 8B model at 4K context and the same model at 64K context are not equivalent memory workloads.
The KV cache grows as the active context grows.
Mistake 3: Planning to the Last Megabyte
If your GPU has 12 GB and your estimated workload requires 11.9 GB, that isn't a comfortable configuration.
Other applications and runtime allocations can push the workload over the limit.
Leave headroom.
How to Reduce VRAM Usage
If you're running into an out-of-memory error, try these approaches in order:
1. Reduce the Context
For example:
-c 8192instead of:
-c 32768This can substantially reduce KV-cache requirements.
2. Use a Smaller Quantization
Moving from a higher-precision model to a 4-bit version can significantly reduce weight memory.
3. Quantize the KV Cache
If your inference engine supports it, consider FP8 or another supported KV-cache format.
4. Reduce Concurrent Requests
Multiple active sequences can require additional KV-cache capacity.
5. Offload Some Layers to CPU
If the model is only slightly too large for your GPU, CPU offloading can make it usable at the cost of performance.
6. Close Other GPU-Heavy Applications
Before benchmarking, close applications that are consuming significant GPU memory.
The Practical Formula
When estimating whether a model will fit, think about it like this:
Available GPU Memory
↓
Model Weights
+
KV Cache
+
Runtime Memory
+
Other GPU Usage
+
Safety Headroom
↓
Should remain below available memoryThat's more reliable than using a single “B parameters = X GB” chart.
Final Takeaway
The biggest mistake in local LLM hardware planning is treating parameter count as the entire memory requirement.
It isn't.
The model weights are only the starting point.
Context length, KV cache precision, runtime behavior, concurrent requests, GPU memory already in use, and CPU offloading can all change the final result.
Before downloading a model, check:
- Actual quantized model size
- Model architecture
- Target context length
- KV-cache precision
- Inference framework
- Available GPU memory
- Whether CPU offloading is acceptable
For a quick estimate, remember:
Model weights tell you whether the model can load.
The KV cache tells you how much context you can afford.
Runtime overhead and available headroom determine whether the whole workload actually fits.
That is the difference between downloading a model that looks compatible on paper and running one reliably in practice.
Sources & References
- llama.cpp documentation — server context configuration, KV-cache options, and GPU offloading. (GitHub)
- vLLM documentation — GPU memory management and KV-cache data types. (vLLM)
- vLLM quantized KV-cache documentation — FP8 KV-cache memory optimization. (vLLM)
- Llama 3.1 8B model configuration — architecture values used in the KV-cache example. (Hugging Face)
- Microsoft Windows graphics memory documentation — GPU memory management under WDDM. (Microsoft Learn)
Editor's Note: VRAM requirements vary between models, quantization formats, inference engines, context settings, and hardware. The calculations in this article are intended for planning and should not be treated as guaranteed runtime measurements. Always check the actual model configuration and inference-engine documentation before deployment.
