LW IT Solutions logo LW IT Solutions logo LW IT Solutions
« Blog Overview /Cloud & AI / Q4, Q8, FP16 and Long Context: Why...
Read this article in other languages:

Q4, Q8, FP16 and Long Context: Why a Local LLM Can Fail Despite Fitting on Disk

Q4, Q8, FP16 and Long Context: Why a Local LLM Can Fail Despite Fitting on Disk
Contents
  1. Q4, Q8 and FP16 imply different budgets
  2. Context consumes runtime memory
  3. Hybrid architectures need their own profiles
  4. Offloading changes the workload
  5. Check quality before optimising memory
  6. Related tools
  7. Questions and answers
  8. Sources

A 20 GB model file does not automatically fit on a 24 GB graphics card. The download describes stored weights and metadata. Execution adds context state, temporary buffers and other allocations. Available capacity also falls below advertised memory when other applications use the GPU.

Quantisation and context therefore belong in the same plan. The following comparison explains memory mechanisms and calculations without deriving measured quality or speed rankings.

More than the model file: Weights; KV / model state; Runtime buffers; Other processes; Free headroom
1 GiB cache in the assumed example; not a universal model formula.

Q4, Q8 and FP16 imply different budgets

Idealised example Calculation for 30 billion weights Weights only
4 bit 30 billion × 0.5 byte 15 GB
8 bit 30 billion × 1 byte 30 GB
16 bit 30 billion × 2 bytes 60 GB

These are theoretical decimal quantities, not file sizes for a particular model. Scales, blocks, mixed precision and unquantised components alter the result. GB and GiB also differ: 15 billion bytes are approximately 13.97 GiB. A Q4 file need not use exactly four bits for each parameter in the model.

Context consumes runtime memory

With conventional attention, the KV cache stores earlier keys and values. A simplified dense-cache calculation is: two × layers × KV heads × head dimension × tokens × bytes per value × sequences. This requires matching architecture and runtime assumptions.

An assumed example with 32 layers, eight KV heads, dimension 128, 8,192 tokens and two-byte values needs 1 GiB for one sequence. Four fully occupied equivalent sequences need 4 GiB. Weights and runtime buffers are additional. Paging, cache quantisation and other attention patterns change actual allocation.

Hybrid architectures need their own profiles

Qwen3.5 uses a hybrid structure with different layer types. A conventional-attention formula cannot be applied blindly to every layer. Recurrent state, attention cache and multimodal components require appropriate model configuration and runtime data.

MoE also separates active computation from total weight storage. A small active-parameter count does not promise a small model file. Planning solely from a model name misses these differences.

Offloading changes the workload

llama.cpp supports hybrid CPU/GPU execution. Partial offloading may make an otherwise oversized model runnable. This remains a different configuration from full GPU execution. System memory, CPU performance and transfers influence response behaviour.

A useful comparison records GPU allocation, system RAM, first-token latency and generation speed together. Successful startup is only the first check. Long documents and concurrent sessions can later expose a limit hidden by a short chat.

Check quality before optimising memory

A larger Q4 model does not automatically outperform a smaller Q8 model in quality. Training, architecture and task remain decisive. A shared dataset with expected answers reveals errors, omissions and format problems; file size alone does not establish usefulness.

The existing memory estimator provides an initial shortlist. Specific revision and backend, free headroom, context and concurrency belong in its inputs. Final configuration follows a real memory check and a small quality evaluation. An unknown architecture needs a visible estimation limit rather than a falsely exact result.

Memory planning for model selection: LLM VRAM estimator.

Questions and answers

Why does a 20 GB model file not always fit on a 24 GB GPU?

Runtime buffers, model state or KV cache and other allocations add to the file size. Actual memory use depends on runtime, architecture, context and concurrency.

Is a larger Q4 model always better than a smaller Q8 model?

Precision alone does not determine a quality ranking. Model family, training and task remain decisive. A comparison requires equal tasks, documented configurations and error assessment rather than file size alone.

Sources

  1. llama.cpp: Projekt und Backends
  2. Qwen: Qwen3.5-9B, offizielle Model Card
  3. Mistral: Ministral 3 14B Instruct, offizielle Model Card
  4. vLLM: GPU installation

Sources checked: 5 October 2026.

Lukas Wojcik

Lukas Wojcik

Systems architect and technology enthusiast specializing in scalable tracking solutions, GMP Stack (GA4 & GTM), and robust backend architectures. Advocate for clean code and privacy-first design.

Get in Touch

Briefly describe your project or inquiry for a tailored response. This site is protected by reCAPTCHA.

Write a comment

Experience with other models or providers and questions about the implementation are welcome here.

The email address is not published. Required fields are marked with an asterisk.

Articles & categories

CCTV

All 6 articles in this category Follow this category by RSS

Cloud & AI

All 19 articles in this category Follow this category by RSS

Data Privacy

All 20 articles in this category Follow this category by RSS

Digital Analytics

All 60 articles in this category Follow this category by RSS

Digital Marketing

All 39 articles in this category Follow this category by RSS

IT & Networks

All 19 articles in this category Follow this category by RSS

Music Production

All 18 articles in this category Follow this category by RSS

Raspberry PI

All 11 articles in this category Follow this category by RSS

SaaS & Internet Earning

Follow this category by RSS

Smart Home

All 18 articles in this category Follow this category by RSS

Web Development

All 11 articles in this category Follow this category by RSS

WordPress Plugins & Tricks

All 15 articles in this category Follow this category by RSS