GGUF reference
GGUF quantization and projected model size
Quantization stores model weights with fewer bits than full 16- or 32-bit values. It reduces file size and weight memory, with a model-dependent tradeoff in output quality.
What a quantization name tells you
Names such as Q4_K_M describe a quantization recipe, not a promise that every tensor uses exactly four bits. Mixed schemes keep some tensors at higher precision and add scale or block data. That is why the inspector works with effective bits-per-weight ranges.
Lower-bit profiles are generally smaller. Higher-bit profiles generally retain more numerical detail. The effect on a specific model depends on its architecture, calibration, task, converter, and runtime, so the tool does not assign quality scores.
Profiles used by the inspector
| Profile | Effective bits per weight used for projection |
|---|---|
| Q2_K | 2.8–3.2 |
| Q3_K_S / M / L | 3.3–3.7 / 3.7–4.1 / 4.1–4.5 |
| Q4_0 | 4.4–4.7 |
| Q4_K_S / M | 4.4–4.8 / 4.7–5.1 |
| Q5_K_S / M | 5.3–5.7 / 5.6–6.0 |
| Q6_K | 6.5–6.9 |
| Q8_0 | 8.3–8.7 |
| F16 / BF16 | 16.0–16.4 |
These are projection profiles in the estimator. The GGUF parser itself preserves each tensor's numeric GGML type code and does not label or validate all possible quantization types.
How projected size is calculated
The tool multiplies the dimensions of every tensor and adds those products to get a parameter count. If a dimension is missing, non-positive, or would make the count unsafe in JavaScript, projections are not shown.
projected file bytes = projected weight bytes + current tensor-data offset
The current effective bits per weight are calculated from the uploaded file's bytes after the tensor-data offset. The profile with the nearest midpoint is marked “Closest to uploaded size.” This is a size comparison, not reliable detection of the file's actual quantization recipe.
Projected total RAM uses the selected context and parallel-sequence settings:
Why a converted file can differ
An actual conversion can choose a different type for individual tensors, preserve output or embedding tensors at higher precision, use updated quantizer rules, and add or change metadata. Alignment and architecture-specific tensor mixes also affect the result. The shown range is therefore an estimate, not an exact output size.
The projection keeps the current file's tensor-data offset as fixed overhead. A real conversion may change that offset. It also keeps the current KV-cache assumption; weight quantization does not necessarily determine cache precision.
For best quality, convert from an F16 or BF16 source when possible. Requantizing a file that is already quantized can compound information loss.
Inspect your GGUF to calculate projections from its own tensor shapes. Read the memory guide before treating projected RAM as a hardware requirement.