GPU Memory Math for Local LLMs: How to Calculate VRAM Before Downloading

Dileep Solanki


Downloading a large local AI model only to discover that your GPU cannot actually run it is an easy mistake to make.

A model may be advertised as an “8B” model, but parameter count alone doesn't tell you how much GPU memory inference will require.

The memory used during inference depends on several things:

  • Model weights
  • Quantization format
  • Context length
  • KV cache
  • Batch or concurrent requests
  • Inference framework
  • GPU memory already being used by other applications

That is why an 8B model can fit comfortably in one configuration and still run out of memory in another.

The useful question isn't:

“How large is the model?”

It's:

“How much memory will this particular model configuration need at my target context length?”

This guide explains how to estimate that before downloading a model.


The Basic Inference Memory Equation

A simplified model of GPU memory usage is:

Total GPU Memory ≈ Model Weights + KV Cache + Runtime/Activation Memory + Other GPU Usage

This is an estimate rather than an exact universal formula.

Different inference engines allocate memory differently, and features such as batching, parallel requests, Flash Attention, quantized KV caches, and CPU offloading can change the final number.

So the goal of the calculation is not to predict the exact number shown by nvidia-smi.

The goal is to determine whether you have enough headroom to run the configuration you actually want.


1. Start With Model Weight Memory

The first component is the model's weights.

A simple lower-bound estimate is:

Weight Memory ≈ Parameter Count × Bits per Parameter ÷ 8

For example, an 8-billion-parameter model stored at approximately 4 bits per parameter would require:

8B × 4 ÷ 8 ≈ 4 GB

But that does not mean every 4-bit 8B model will occupy exactly 4 GB.

Quantized model files contain additional information, and different quantization schemes have different overheads.

The actual model file size is therefore a better starting point than multiplying the parameter count by a single fixed multiplier.


Approximate Weight-Only Estimates

A useful first-pass calculation looks like this:

ModelApproximate raw weight calculation
7B at FP16~14 GB
8B at FP16~16 GB
14B at FP16~28 GB
32B at FP16~64 GB
70B at FP16~140 GB

For lower-bit quantization, the theoretical weight requirement falls substantially.

For example:

Model8-bit4-bit theoretical minimum
8B~8 GB~4 GB
14B~14 GB~7 GB
32B~32 GB~16 GB
70B~70 GB~35 GB

These are planning estimates, not exact download sizes or guaranteed VRAM requirements.

For GGUF models in particular, check the actual file size and quantization type before downloading.


2. The Part Many People Forget: KV Cache

Model weights are relatively static.

The KV cache is different.

During generation, the model stores attention information from previously processed tokens so it doesn't have to recompute the same information repeatedly.

As the context grows, the KV cache grows too.

This is why a model that works at 4K context can run out of memory at 32K or 128K context.


How to Calculate KV Cache Memory

For a standard transformer using grouped-query attention, a useful simplified formula is:

KV Cache per Token = 2 × Layers × KV Heads × Head Dimension × Bytes per Element

The 2 represents the Key and Value tensors.

For example, Llama 3.1 8B has:

  • 32 transformer layers
  • 8 key/value heads
  • 128-dimensional attention heads

These architecture values are documented in the model configuration. (Hugging Face)

Assuming FP16 KV storage, where each value uses 2 bytes:

2 × 32 × 8 × 128 × 2

= 131,072 bytes per token

That's approximately:

128 KiB per token

This gives us a useful illustration of how context length affects memory.


Llama 3.1 8B KV Cache Example

Using the simplified FP16 calculation above:

ContextApproximate KV Cache
2K tokens~256 MiB
8K tokens~1 GiB
32K tokens~4 GiB
128K tokens~16 GiB

The important point is the growth.

At 8K context, the KV cache in this simplified example is around 1 GiB.

At 32K, it is around 4 GiB.

At 128K, it reaches around 16 GiB.

That means a model can fit comfortably when running a short context and still run out of memory when you start feeding it long documents.

Actual memory usage can differ because inference engines may use different KV layouts, cache precisions, attention implementations, batching strategies, and other optimizations.


Why Context Length Matters So Much

Consider an 8B model whose quantized weights occupy roughly 5 GB.

You might have a 12 GB GPU and conclude:

12 GB - 5 GB = 7 GB remaining.

That looks comfortable.

But the remaining memory also needs to accommodate the KV cache and runtime allocations.

At a sufficiently long context, the KV cache can consume several additional gigabytes.

And if another application is already using GPU memory, your available headroom becomes even smaller.

This is why model size alone isn't enough to determine whether an LLM will fit.


3. Runtime Memory Also Matters

The remaining GPU memory isn't simply “free space.”

Your inference framework needs memory for additional operations and data structures.

Depending on the framework and workload, this can include:

  • Temporary tensors
  • Attention operations
  • CUDA allocations
  • Model buffers
  • CUDA graphs or related runtime structures
  • Batch processing
  • Multiple concurrent sequences
  • Framework-specific memory management

The exact amount varies considerably.

For example, vLLM exposes GPU-memory controls such as gpu_memory_utilization and has its own KV-cache memory management. Its current documentation lists a default gpu_memory_utilization value of 0.92. (vLLM)

This is one reason a universal “add exactly 1 GB” rule should not be treated as a precise engineering calculation.


4. Your Desktop May Already Be Using VRAM

If you're running a local model on a desktop GPU, the operating system and other applications may already be using some GPU memory.

Your browser, games, video applications, desktop compositor, IDE, and other GPU-accelerated applications can all consume resources.

Windows also manages dedicated and system GPU memory through its graphics memory manager rather than treating VRAM as a completely isolated pool available to one application. (Microsoft Learn)

So if your GPU has:

12 GB VRAM

you shouldn't assume that all 12 GB will be available to the inference engine.

The amount available depends on what else is running.


A Better Practical Formula

Instead of claiming an exact universal formula, use this as a planning model:

Required GPU Memory ≈ Weight Memory + KV Cache + Runtime Overhead + Safety Headroom

For example:

Model weights       ≈ 5 GB
KV cache             ≈ 2 GB
Runtime overhead     ≈ 1 GB+
Safety headroom      ≈ 1–2 GB
--------------------------------
Planning requirement ≈ 9–10 GB+

This doesn't guarantee that the model will run.

It tells you that a 12 GB GPU is in a more comfortable position than an 8 GB GPU for that particular configuration.


A Worked Example: 8B Model on a 12 GB GPU

Suppose you have:

GPU: 12 GB VRAM
Model: 8B
Quantization: approximately 4-bit
Target context: 8K

Start with the model weights.

Assume the downloaded quantized model occupies roughly:

~5 GB

Then estimate the KV cache.

For the Llama 3.1 8B example at FP16 KV precision:

8K context ≈ 1 GiB of KV cache

Now you still need memory for runtime allocations and whatever your GPU is already using.

A rough planning model might therefore look like:

5 GB     Model
~1 GiB   KV cache
+        Runtime allocations
+        Existing GPU usage
+        Safety margin
---------------------------

This could fit on a 12 GB GPU, but whether it actually does depends on the inference engine, operating environment, model file, context configuration, and other GPU workloads.

That's much more useful than simply saying:

“8B models need 8 GB.”

They don't.


Quantization Is Only Part of the Equation

Quantization reduces the memory required for model weights, but it doesn't eliminate the KV cache.

For example:

Higher precision model
        ↓
Larger weights
        ↓
More VRAM

Lower precision model
        ↓
Smaller weights
        ↓
More VRAM available for context

This is one reason quantized models are so useful for local inference.

However, if your context becomes very large, the KV cache can become a significant part of the total memory requirement even when the model weights are heavily quantized.


KV Cache Quantization

Some inference engines can also reduce the memory used by the KV cache.

Current llama.cpp supports configurable KV-cache types through options such as:

-ctk q8_0
-ctv q8_0

for the Key and Value cache respectively. Its current server documentation also supports several other KV-cache types. (GitHub)

For example:

llama-server \
  -m model.gguf \
  -ctk q8_0 \
  -ctv q8_0 \
  -c 16384

This can reduce KV-cache memory compared with FP16 storage.

The exact memory savings and quality impact depend on the model and workload, so treat KV quantization as an optimization to benchmark rather than assuming that it is completely lossless.


vLLM Also Supports KV Cache Quantization

vLLM currently supports several KV-cache data types, including FP8 variants. (vLLM)

For example:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --kv-cache-dtype fp8

vLLM documents FP8 KV-cache quantization as a way to reduce KV-cache memory and support larger token capacity. (vLLM)

Again, compatibility depends on your hardware, model, backend, and vLLM version.


5. Set a Realistic Context Window

One of the easiest ways to control memory usage is to avoid requesting a context window larger than your application actually needs.

With llama-server, the -c / --ctx-size option controls the prompt context size. The current documentation also notes that the default can come from the model configuration when set to zero. (GitHub)

For example:

llama-server \
  -m model.gguf \
  -c 8192

If your application only needs approximately 8K tokens, there's little reason to configure it for 128K simply because the model supports that maximum.

Remember:

Maximum supported context ≠ context you should always allocate.


6. CPU Offloading Can Help

If the complete model doesn't fit into GPU memory, some inference engines can place part of the workload in system RAM.

With llama.cpp, for example, GPU layer offloading can be controlled with options such as -ngl. (GitHub)

A simplified example is:

llama-server \
  -m model.gguf \
  -ngl 24

This doesn't magically make the model as fast as a fully GPU-resident model.

Instead, it trades performance for the ability to run a model that would otherwise exceed the available GPU memory.

The exact number of layers that fit depends on the model, quantization, GPU, context configuration, and other allocations.


What Different VRAM Sizes Actually Mean

Instead of assigning one fixed model size to each GPU, think in terms of workload ranges.

8 GB VRAM

A practical target is generally smaller quantized models.

You can experiment with:

  • 3B–8B models
  • 4-bit quantization
  • Moderate context lengths
  • CPU offloading when necessary

An 8B 4-bit model may be possible, but you need to account for KV cache and runtime memory.

Don't assume that the entire 8 GB is available to the model.


12 GB VRAM

This gives you more flexibility for local development.

Typical possibilities include:

  • 7B–8B models at higher quantization levels
  • 12B–14B models at lower quantization levels
  • Moderate context windows
  • Some CPU/GPU hybrid configurations

Whether a particular 14B model fits depends on its actual quantized file size and the context you configure.


16 GB VRAM

16 GB opens up more options for local experimentation.

You can consider:

  • 7B–14B models at relatively high precision
  • Some 30B-class models with aggressive quantization
  • Longer context configurations for smaller models
  • More room for runtime overhead

But again, the model's actual architecture and quantization matter more than the parameter count alone.


24 GB VRAM

24 GB is a useful capacity for serious local experimentation.

Depending on the model and quantization, you can explore:

  • 14B-class models with relatively high precision
  • 30B–35B-class models using lower-bit quantization
  • Longer context configurations
  • Larger models using partial CPU offloading

A 70B model can still be far beyond the comfortable capacity of a single 24 GB GPU, especially at higher precision.


Apple Silicon Is Different

Apple Silicon systems use unified memory rather than a conventional discrete GPU VRAM pool.

That means the CPU and GPU share system memory.

A Mac with 64 GB of unified memory therefore isn't directly comparable to a desktop GPU with 64 GB of dedicated VRAM.

The amount available to a model depends on the operating system, application, model runtime, and other processes.

For Apple Silicon users, the practical question is therefore:

How much unified memory remains available for the inference workload?

rather than simply:

How much VRAM does my GPU have?


A Simple Pre-Download Calculation

Before downloading a model, collect these numbers:

1. Model size

Check the actual quantized model file.

For example:

Q4 model
≈ 5 GB

2. Context length

Decide what you actually need:

4K
8K
16K
32K

Don't automatically select the model's maximum context.

3. KV cache

Estimate KV memory using the model architecture:

KV Bytes/Token = 2 × Layers × KV Heads × Head Dimension × Bytes/Element

Then multiply by the number of tokens you want to support.

4. Runtime headroom

Leave space for:

  • Inference engine allocations
  • CUDA/runtime memory
  • Other GPU applications
  • Batch processing
  • Multiple concurrent requests

5. Safety margin

Don't plan for a configuration that uses essentially 100% of your GPU memory.

A little headroom makes troubleshooting considerably easier.


A Useful VRAM Planning Table

GPU MemoryPractical Starting Point
8 GBSmall-to-medium quantized models
12 GB7B–14B quantized models
16 GBLarger 14B-class models and some aggressively quantized larger models
24 GB14B–35B-class models depending on quantization and context
32 GB+More room for larger models, longer contexts, and higher precision

These are planning ranges, not hard limits.

A 12 GB GPU can run a model that a generic chart says requires more memory if some layers are offloaded to system RAM.

Likewise, a model that theoretically fits inside 12 GB can still fail if the context window or runtime allocations push total usage beyond the available memory.


Three Common VRAM Mistakes

Mistake 1: Looking Only at the Download Size

A 5 GB model file does not mean the application will use exactly 5 GB of VRAM.

The runtime needs additional memory.


Mistake 2: Ignoring Context Length

An 8B model at 4K context and the same model at 64K context are not equivalent memory workloads.

The KV cache grows as the active context grows.


Mistake 3: Planning to the Last Megabyte

If your GPU has 12 GB and your estimated workload requires 11.9 GB, that isn't a comfortable configuration.

Other applications and runtime allocations can push the workload over the limit.

Leave headroom.


How to Reduce VRAM Usage

If you're running into an out-of-memory error, try these approaches in order:

1. Reduce the Context

For example:

-c 8192

instead of:

-c 32768

This can substantially reduce KV-cache requirements.

2. Use a Smaller Quantization

Moving from a higher-precision model to a 4-bit version can significantly reduce weight memory.

3. Quantize the KV Cache

If your inference engine supports it, consider FP8 or another supported KV-cache format.

4. Reduce Concurrent Requests

Multiple active sequences can require additional KV-cache capacity.

5. Offload Some Layers to CPU

If the model is only slightly too large for your GPU, CPU offloading can make it usable at the cost of performance.

6. Close Other GPU-Heavy Applications

Before benchmarking, close applications that are consuming significant GPU memory.


The Practical Formula

When estimating whether a model will fit, think about it like this:

Available GPU Memory

        ↓

Model Weights
        +
KV Cache
        +
Runtime Memory
        +
Other GPU Usage
        +
Safety Headroom

        ↓

Should remain below available memory

That's more reliable than using a single “B parameters = X GB” chart.


Final Takeaway

The biggest mistake in local LLM hardware planning is treating parameter count as the entire memory requirement.

It isn't.

The model weights are only the starting point.

Context length, KV cache precision, runtime behavior, concurrent requests, GPU memory already in use, and CPU offloading can all change the final result.

Before downloading a model, check:

  1. Actual quantized model size
  2. Model architecture
  3. Target context length
  4. KV-cache precision
  5. Inference framework
  6. Available GPU memory
  7. Whether CPU offloading is acceptable

For a quick estimate, remember:

Model weights tell you whether the model can load.

The KV cache tells you how much context you can afford.

Runtime overhead and available headroom determine whether the whole workload actually fits.

That is the difference between downloading a model that looks compatible on paper and running one reliably in practice.

Sources & References

  • llama.cpp documentation — server context configuration, KV-cache options, and GPU offloading. (GitHub)
  • vLLM documentation — GPU memory management and KV-cache data types. (vLLM)
  • vLLM quantized KV-cache documentation — FP8 KV-cache memory optimization. (vLLM)
  • Llama 3.1 8B model configuration — architecture values used in the KV-cache example. (Hugging Face)
  • Microsoft Windows graphics memory documentation — GPU memory management under WDDM. (Microsoft Learn)

Editor's Note: VRAM requirements vary between models, quantization formats, inference engines, context settings, and hardware. The calculations in this article are intended for planning and should not be treated as guaranteed runtime measurements. Always check the actual model configuration and inference-engine documentation before deployment.

3/related/default