Calculation reference

How the inference memory estimate works

The report estimates CPU memory for one loaded model. It adds model weights, a KV-cache estimate when the file exposes enough information, and a fixed rule for runtime buffers.

The total

estimated total = weight bytes + KV-cache bytes + runtime allowance
runtime allowance = max(weight bytes × 10%, 256 MiB)

The allowance is a simple safety margin in the application, not a measurement of a particular runtime. The “weights + cache” card excludes it; the “estimated total” card includes it.

GGUF weight and KV-cache formulas

GGUF weight bytes are the uploaded file size minus the aligned byte offset where tensor data begins. This includes the whole tensor-data region and any bytes after its start; the tool does not sum encoded tensor blocks.

GGUF weight bytes = file size − tensor-data offset

For the cache, the estimator uses architecture-prefixed metadata:

KV bytes = blocks × KV heads × (key length + value length) × 2 bytes × context tokens × parallel sequences

Blocks is {architecture}.block_count. KV heads is attention.head_count_kv, falling back to attention.head_count. Key and value length come from explicit metadata when available. Otherwise key length is embedding_length / head_count, and value length falls back to key length. The 2-byte factor assumes 16-bit key and value storage.

Context defaults to the architecture's declared context_length. If it is absent, the default is 4,096. Parallel sequences default to one. Cache memory grows linearly with both controls: doubling context or parallel sequences doubles this part of the estimate.

Worked GGUF example

Consider a file of 7,000,000,000 bytes whose tensor data begins at byte 4,000,000. Its metadata reports 40 blocks, 40 attention heads, 10 KV heads, and an embedding length of 5,120. With no explicit key length, the tool derives 5,120 / 40 = 128. Value length also becomes 128. At 16,384 tokens and one sequence:

weights = 7,000,000,000 − 4,000,000 = 6,996,000,000 bytes
KV cache = 40 × 10 × (128 + 128) × 2 × 16,384 × 1
KV cache = 3,355,443,200 bytes
allowance = 699,600,000 bytes
total = 11,051,043,200 bytes (about 10.29 GiB)

This example matches the estimator test fixture. It is not a recommendation for a particular model or runtime.

ONNX weights and cache

For ONNX, the tool multiplies each initializer's dimensions and its data-type byte width, then adds the tensors it can count. It recognizes standard numeric types represented by the parser, including half, bfloat16, float8, and 4-bit types. The calculation includes initializers whose tensor content is stored as external data because it uses the declared shape and type, not the bytes embedded in the main ONNX file.

The parser does not currently read external-data locations or verify those files. It also ignores sparse-initializer payloads, initializers with missing dimensions, and types without a known byte width.

An ONNX KV cache is counted only when graph inputs or outputs have names matching past_key_values...key/value or present...key/value. Outputs are preferred when present. A symbolic dimension containing “batch” uses parallel sequences; one containing “sequence” or “context” uses context length. Other unresolved dimensions are treated as one. The tensor is counted only if a sequence/context dimension was identified.

What the estimate leaves out

Actual memory can differ because runtimes use memory mapping, temporary workspaces, graph optimizations, allocator padding, different cache precision, quantized caches, or CPU/GPU offloading. Some architectures store additional state or use attention arrangements that the metadata formula does not describe. The estimate also does not model operating-system memory, a tokenizer, application code, or multiple loaded models.

For GGUF, a missing required field produces a zero KV-cache estimate and a medium-confidence label. For ONNX, “high” confidence only means every initializer was counted, explicit cache tensors were found, and their non-batch dimensions were resolved. It does not mean measured accuracy against a runtime.

Treat the number as a planning estimate and leave headroom. Use the inspector to change context and parallel sequences for your file.