Skip to main content
Home/Tools/Developer/Local AI Deployment Lab

Local AI Deployment Lab

Plan a private LLM deployment from model and hardware fit through estimated speed and reviewable Ollama or llama.cpp output.

100% Private - Runs Entirely in Your Browser
No data is sent to any server. All processing happens locally on your device.

What Context Length Costs: num_ctx, KV Cache, and Tokens Per Second

Raising a local model’s context window is not free, and the price is charged twice. The KV cache grows linearly with every token of context you allow, taking memory away from the weights, and because generation speed is bounded by how much memory has to be read per token, a larger cache also makes the model slower. This tool shows both costs at once: pick a model, a quantization, a context length, and hardware, and it splits the memory budget into weights, KV cache, and runtime overhead, then estimates the generation speed that combination produces.

It is the answer to the question that comes up the moment a local model is working — “can I raise num_ctx to 32k?” — and to the follow-up nobody asks first, which is what that does to throughput. Nothing is downloaded and no model is executed; this is a planning model, and the output includes the runtime configuration to go and try it.

Why the KV Cache Is Model-Specific

There is no universal “GB per thousand tokens” figure, and rules of thumb quoted from one model are badly wrong applied to another. The cache is computed from the model’s own attention geometry:

KV bytes = 2 × KV heads × head dimension × bytes per element × layers × context

Three things in that expression vary enormously between models of the same nominal size:

  • KV heads, not attention heads. Grouped-query attention lets many query heads share one key-value head. A model with 32 attention heads but 8 KV heads has a quarter of the cache a multi-head model of identical parameter count would need. This is the single biggest source of variation.
  • Sliding-window layers. Several modern architectures alternate global-attention layers with local ones whose cache is capped at the window size. Those layers stop growing once context exceeds the window, so the total cache grows more slowly than linearly — the model is calculated layer by layer with global and local layers counted separately.
  • Latent attention. Models using multi-head latent attention store one compressed vector per token per layer, shared across heads, which changes the arithmetic entirely rather than merely scaling it.

On top of that, KV cache precision is selectable. FP16 uses two bytes per element; Q8 uses one and halves the cache outright. When a configuration misses the memory budget by a little, switching KV precision is usually a smaller quality compromise than dropping the weights to a more aggressive quantization — and it is the first thing to try before giving up context.

The Full Memory Budget

Three components are modelled and shown as a stacked bar, so you can see which one is actually consuming the budget:

ComponentHow it is calculatedScales with
WeightsParameters × bytes per parameter for the chosen quantizationModel size and quantization only
KV cacheAttention geometry × context length × KV precisionContext length, linearly (less for sliding-window layers)
Runtime overhead6% of weights, with a 0.75 GiB floorModel size

Against that total sits usable memory, and the distinction there matters. A discrete GPU contributes its full VRAM. Unified-memory systems and CPU inference are counted at 75% of installed RAM, because the operating system and everything else running need the rest — a 64 GB Apple Silicon machine is modelled as 48 GB of usable budget, not 64. Multiple units are summed only for hardware profiles that actually support being combined.

A verdict is shown outright, with headroom below 10% raised as a separate finding even when the configuration technically fits. That warning is worth heeding: real runtimes allocate beyond the modelled overhead, prompt length varies, and a configuration sitting at 95% of budget is one long prompt away from an out-of-memory failure.

Context Length Also Costs Speed

Single-user generation on local hardware is bounded by memory bandwidth, not compute. Each token requires reading the active weights and the KV cache, so the estimate is:

tokens/sec ≈ memory bandwidth × 0.625 / (active weights + KV cache)

The KV cache being in the denominator is the part people miss. Raising context does not merely consume memory that was sitting idle — it directly reduces generation speed, and the effect grows as the cache becomes a larger share of what must be read per token. On a small model with a large context, the cache can dominate the weights entirely.

The 0.625 factor is an efficiency allowance covering the gap between theoretical bandwidth and what a real runtime achieves. For mixture-of-experts models only the active parameters count, which is why they generate far faster than their total parameter count suggests. Tensor-parallel splits across multiple units scale at 90% per additional unit; layer splits do not scale throughput at all, because layers execute in sequence — layer splitting buys capacity, not speed, and expecting two GPUs to double tokens per second in that mode is a common disappointment.

Treat the result as a bounded estimate, not a benchmark. Batching, runtime kernels, prompt length, thermal throttling, and any offload to system memory all move real numbers. To measure rather than model, the LLM GPU benchmark runs on your actual hardware.

Runtime Configuration You Can Actually Run

Because a context length that only exists in a planning tool helps nobody, the chosen configuration is emitted as runnable output for either runtime.

For Ollama, a Modelfile that sets the context explicitly and the two commands to build and run it. This is the step that is easy to skip: Ollama applies its own default context window rather than the model’s published maximum, so a model advertising 128k tokens will quietly truncate at a much smaller window until num_ctx is set. If a local model appears to forget the start of a long document, the context window is the first thing to check, and it is usually the configuration rather than the model.

For llama.cpp, a llama-cli invocation with -c set to the context length and -ngl set to offload all layers to the GPU — or to zero when the profile is CPU-only. Partial offload, where some layers sit in VRAM and the rest in system RAM, runs at the speed of the slowest path and is not modelled here; if a configuration does not fit, reduce context or quantization rather than relying on spillover.

Context length is validated against the model’s published maximum, so a value the model cannot support is rejected with the real limit rather than silently accepted. The range is 256 to 1,048,576 tokens, with 1 to 16 hardware units.

What Gets Flagged

  • Does not fit — high. The suggested moves, in rough order of quality cost: reduce context, switch KV cache to Q8, choose a more aggressive weight quantization, or pick a smaller model.
  • Headroom below 10% — medium. It fits on paper and may not fit in practice.
  • Aggressive quantization selected — medium, when the dataset marks the chosen quantization as low quality. Validate task quality before committing.
  • CPU inference selected — low. Capacity is often fine; interactive speed usually is not.
  • Hardware does not combine multiple units — low, when a unit count above one is entered for a profile that cannot pool memory. The calculation falls back to a single unit.

The model and hardware dataset carries a visible as-of date, shown on the results panel, so you can judge how current the figures are rather than assuming.

How to Size a Context Window

  1. Choose the model and quantization you intend to run. The panel shows its Hugging Face identifier and published maximum context.
  2. Set the context length you actually need — the real length of your documents and conversations, not the largest number the model accepts.
  3. Choose your hardware and unit count, and the parallel mode if more than one.
  4. Read the stacked memory bar. If KV cache rivals or exceeds the weights, context is what is costing you.
  5. Check the generation estimate alongside the fit verdict — a configuration that fits but generates at a few tokens per second is not usable interactively.
  6. If it does not fit, try Q8 KV cache first, then reduce context, then quantization, then model size.
  7. Copy the Modelfile or llama.cpp command and run the configuration you just sized.

Frequently Asked Questions

How much VRAM does context length use in Ollama?

It depends on the model’s attention geometry, not just its size, so there is no single figure. The cache is two bytes per element (or one at Q8) times the KV head count, head dimension, and layer count, multiplied by the context length. Two models of the same parameter count can differ several-fold. Select your model above and the exact figure is shown for the context you choose.

Why does my model truncate long documents even though it supports 128k context?

Almost always because the runtime is applying its own default context window rather than the model’s maximum. In Ollama, set num_ctx in the Modelfile; in llama.cpp, pass -c. The Modelfile and command emitted above set it explicitly. Then check the result fits — a 128k window on a model whose cache is expensive may not.

Does a longer context make generation slower?

Yes, and this is the cost people overlook. Generation speed is bounded by memory read per token, and the KV cache is part of what must be read. Raising context both consumes memory and reduces tokens per second, with the effect growing as the cache becomes a larger share of the total.

Should I lower the KV cache to Q8?

It is usually the first thing to try when a configuration is close to fitting. It halves the cache outright, and the quality impact is generally smaller than dropping the weights a quantization level. Validate on your own task rather than assuming either way.

Will two GPUs double my tokens per second?

Only with tensor parallelism, and even then the model allows about 90% scaling per additional unit. A layer split runs layers in sequence across devices, so it buys capacity for a model that would not otherwise fit, not throughput. Check which mode your runtime is actually using before expecting a speedup.

Why is my Apple Silicon machine’s usable memory shown as less than its RAM?

Unified-memory systems and CPU inference are counted at 75% of installed RAM, because the operating system and other applications need the remainder. A 64 GB machine is modelled with a 48 GB budget. Discrete GPUs contribute their full VRAM.

Are these numbers measured?

No. Memory is computed from published model geometry and the speed figure is a bandwidth-bound estimate at a fixed efficiency factor. It is deliberately a bounded midpoint rather than a benchmark. For measured numbers on your own hardware, run the GPU benchmark.

Do I need an account?

No. This and the rest of our developer tools are free and need no signup, and no model is downloaded or executed. For full VRAM exploration across models use the LLM VRAM calculator; to build complete Ollama commands and Modelfiles use the Ollama command builder.

What Context Length Costs: num_ctx, KV Cache, and Tokens Per Second

Raising a local model’s context window is not free, and the price is charged twice. The KV cache grows linearly with every token of context you allow, taking memory away from the weights, and because generation speed is bounded by how much memory has to be read per token, a larger cache also makes the model slower. This tool shows both costs at once: pick a model, a quantization, a context length, and hardware, and it splits the memory budget into weights, KV cache, and runtime overhead, then estimates the generation speed that combination produces.

It is the answer to the question that comes up the moment a local model is working — “can I raise num_ctx to 32k?” — and to the follow-up nobody asks first, which is what that does to throughput. Nothing is downloaded and no model is executed; this is a planning model, and the output includes the runtime configuration to go and try it.

Why the KV Cache Is Model-Specific

There is no universal “GB per thousand tokens” figure, and rules of thumb quoted from one model are badly wrong applied to another. The cache is computed from the model’s own attention geometry:

KV bytes = 2 × KV heads × head dimension × bytes per element × layers × context

Three things in that expression vary enormously between models of the same nominal size:

  • KV heads, not attention heads. Grouped-query attention lets many query heads share one key-value head. A model with 32 attention heads but 8 KV heads has a quarter of the cache a multi-head model of identical parameter count would need. This is the single biggest source of variation.
  • Sliding-window layers. Several modern architectures alternate global-attention layers with local ones whose cache is capped at the window size. Those layers stop growing once context exceeds the window, so the total cache grows more slowly than linearly — the model is calculated layer by layer with global and local layers counted separately.
  • Latent attention. Models using multi-head latent attention store one compressed vector per token per layer, shared across heads, which changes the arithmetic entirely rather than merely scaling it.

On top of that, KV cache precision is selectable. FP16 uses two bytes per element; Q8 uses one and halves the cache outright. When a configuration misses the memory budget by a little, switching KV precision is usually a smaller quality compromise than dropping the weights to a more aggressive quantization — and it is the first thing to try before giving up context.

The Full Memory Budget

Three components are modelled and shown as a stacked bar, so you can see which one is actually consuming the budget:

ComponentHow it is calculatedScales with
WeightsParameters × bytes per parameter for the chosen quantizationModel size and quantization only
KV cacheAttention geometry × context length × KV precisionContext length, linearly (less for sliding-window layers)
Runtime overhead6% of weights, with a 0.75 GiB floorModel size

Against that total sits usable memory, and the distinction there matters. A discrete GPU contributes its full VRAM. Unified-memory systems and CPU inference are counted at 75% of installed RAM, because the operating system and everything else running need the rest — a 64 GB Apple Silicon machine is modelled as 48 GB of usable budget, not 64. Multiple units are summed only for hardware profiles that actually support being combined.

A verdict is shown outright, with headroom below 10% raised as a separate finding even when the configuration technically fits. That warning is worth heeding: real runtimes allocate beyond the modelled overhead, prompt length varies, and a configuration sitting at 95% of budget is one long prompt away from an out-of-memory failure.

Context Length Also Costs Speed

Single-user generation on local hardware is bounded by memory bandwidth, not compute. Each token requires reading the active weights and the KV cache, so the estimate is:

tokens/sec ≈ memory bandwidth × 0.625 / (active weights + KV cache)

The KV cache being in the denominator is the part people miss. Raising context does not merely consume memory that was sitting idle — it directly reduces generation speed, and the effect grows as the cache becomes a larger share of what must be read per token. On a small model with a large context, the cache can dominate the weights entirely.

The 0.625 factor is an efficiency allowance covering the gap between theoretical bandwidth and what a real runtime achieves. For mixture-of-experts models only the active parameters count, which is why they generate far faster than their total parameter count suggests. Tensor-parallel splits across multiple units scale at 90% per additional unit; layer splits do not scale throughput at all, because layers execute in sequence — layer splitting buys capacity, not speed, and expecting two GPUs to double tokens per second in that mode is a common disappointment.

Treat the result as a bounded estimate, not a benchmark. Batching, runtime kernels, prompt length, thermal throttling, and any offload to system memory all move real numbers. To measure rather than model, the LLM GPU benchmark runs on your actual hardware.

Runtime Configuration You Can Actually Run

Because a context length that only exists in a planning tool helps nobody, the chosen configuration is emitted as runnable output for either runtime.

For Ollama, a Modelfile that sets the context explicitly and the two commands to build and run it. This is the step that is easy to skip: Ollama applies its own default context window rather than the model’s published maximum, so a model advertising 128k tokens will quietly truncate at a much smaller window until num_ctx is set. If a local model appears to forget the start of a long document, the context window is the first thing to check, and it is usually the configuration rather than the model.

For llama.cpp, a llama-cli invocation with -c set to the context length and -ngl set to offload all layers to the GPU — or to zero when the profile is CPU-only. Partial offload, where some layers sit in VRAM and the rest in system RAM, runs at the speed of the slowest path and is not modelled here; if a configuration does not fit, reduce context or quantization rather than relying on spillover.

Context length is validated against the model’s published maximum, so a value the model cannot support is rejected with the real limit rather than silently accepted. The range is 256 to 1,048,576 tokens, with 1 to 16 hardware units.

What Gets Flagged

  • Does not fit — high. The suggested moves, in rough order of quality cost: reduce context, switch KV cache to Q8, choose a more aggressive weight quantization, or pick a smaller model.
  • Headroom below 10% — medium. It fits on paper and may not fit in practice.
  • Aggressive quantization selected — medium, when the dataset marks the chosen quantization as low quality. Validate task quality before committing.
  • CPU inference selected — low. Capacity is often fine; interactive speed usually is not.
  • Hardware does not combine multiple units — low, when a unit count above one is entered for a profile that cannot pool memory. The calculation falls back to a single unit.

The model and hardware dataset carries a visible as-of date, shown on the results panel, so you can judge how current the figures are rather than assuming.

How to Size a Context Window

  1. Choose the model and quantization you intend to run. The panel shows its Hugging Face identifier and published maximum context.
  2. Set the context length you actually need — the real length of your documents and conversations, not the largest number the model accepts.
  3. Choose your hardware and unit count, and the parallel mode if more than one.
  4. Read the stacked memory bar. If KV cache rivals or exceeds the weights, context is what is costing you.
  5. Check the generation estimate alongside the fit verdict — a configuration that fits but generates at a few tokens per second is not usable interactively.
  6. If it does not fit, try Q8 KV cache first, then reduce context, then quantization, then model size.
  7. Copy the Modelfile or llama.cpp command and run the configuration you just sized.

Frequently Asked Questions

How much VRAM does context length use in Ollama?

It depends on the model’s attention geometry, not just its size, so there is no single figure. The cache is two bytes per element (or one at Q8) times the KV head count, head dimension, and layer count, multiplied by the context length. Two models of the same parameter count can differ several-fold. Select your model above and the exact figure is shown for the context you choose.

Why does my model truncate long documents even though it supports 128k context?

Almost always because the runtime is applying its own default context window rather than the model’s maximum. In Ollama, set num_ctx in the Modelfile; in llama.cpp, pass -c. The Modelfile and command emitted above set it explicitly. Then check the result fits — a 128k window on a model whose cache is expensive may not.

Does a longer context make generation slower?

Yes, and this is the cost people overlook. Generation speed is bounded by memory read per token, and the KV cache is part of what must be read. Raising context both consumes memory and reduces tokens per second, with the effect growing as the cache becomes a larger share of the total.

Should I lower the KV cache to Q8?

It is usually the first thing to try when a configuration is close to fitting. It halves the cache outright, and the quality impact is generally smaller than dropping the weights a quantization level. Validate on your own task rather than assuming either way.

Will two GPUs double my tokens per second?

Only with tensor parallelism, and even then the model allows about 90% scaling per additional unit. A layer split runs layers in sequence across devices, so it buys capacity for a model that would not otherwise fit, not throughput. Check which mode your runtime is actually using before expecting a speedup.

Why is my Apple Silicon machine’s usable memory shown as less than its RAM?

Unified-memory systems and CPU inference are counted at 75% of installed RAM, because the operating system and other applications need the remainder. A 64 GB machine is modelled with a 48 GB budget. Discrete GPUs contribute their full VRAM.

Are these numbers measured?

No. Memory is computed from published model geometry and the speed figure is a bandwidth-bound estimate at a fixed efficiency factor. It is deliberately a bounded midpoint rather than a benchmark. For measured numbers on your own hardware, run the GPU benchmark.

Do I need an account?

No. This and the rest of our developer tools are free and need no signup, and no model is downloaded or executed. For full VRAM exploration across models use the LLM VRAM calculator; to build complete Ollama commands and Modelfiles use the Ollama command builder.

Loading interactive tool...

You build the idea. I'll ship the product.

Productized MVP development for founders. 9 SaaS apps shipped — yours could be next, in 6 weeks. Secure by default.

Frequently Asked Questions

Common questions about the Local AI Deployment Lab

No. It performs local arithmetic from curated architecture and hardware datasets and generates reviewable starter text. Downloads, commands, servers, and models run only if you choose to use that output elsewhere.

Memory uses model weights, architecture-aware KV cache, and approximately 6% runtime overhead with a minimum allowance. Speed is a single-user bandwidth model. Runtime kernels, batching, offload, thermals, and prompts change real results.

Published bandwidth describes hardware potential, not your complete system. A measured benchmark captures browser and driver support, power limits, memory behavior, and other machine-specific constraints that an estimate cannot observe.

ℹ️ Disclaimer

This tool is provided for informational and educational purposes only. All processing happens entirely in your browser - no data is sent to or stored on our servers. While we strive for accuracy, we make no warranties about the completeness or reliability of results. Use at your own discretion.