GGUF reference

Reading GGUF metadata

GGUF metadata is a typed set of key-value pairs stored before the model tensors. AI Model Inspector displays every pair it can parse, but it interprets only a small group for calculations.

Architecture and naming

general.architecture gives the prefix used for architecture-specific keys. If its value is llama, for example, the estimator looks for keys such as llama.block_count. This matters because different model families can use different attention layouts.

Other general.* fields may record a model name, description, author, license, or quantization version. The metadata table displays such fields, but the current parser does not turn them into assumptions about model quality or compatibility.

Context and model shape

Key patternMeaning in the toolWhy it matters
{architecture}.context_lengthThe declared maximum used as the default and upper context option.Context length scales the GGUF KV-cache estimate.
{architecture}.embedding_lengthThe hidden-state width. It supplies a fallback attention head dimension.Together with head count, it can derive key and value length.
{architecture}.block_countThe number of transformer blocks used in the KV formula.More blocks produce more cached key and value elements.

If the architecture or declared context key is missing, the interface defaults to 4,096 tokens and offers presets up to 32,768. That fallback is a workload assumption, not a claim about the model's real limit.

Attention fields

{architecture}.attention.head_count is the total attention-head count. attention.head_count_kv is the number of key/value heads. When the KV-head count is absent, the estimator uses the total head count.

The estimator first looks for attention.key_length and attention.value_length. If key length is missing, it uses embedding_length / head_count. Missing value length falls back to key length. If block count, KV-head count, key length, or value length still cannot be derived, the displayed KV cache is zero and a note says it could not be calculated.

Tokenizer metadata

GGUF commonly stores tokenizer model information, tokens, scores, token types, and identifiers for special tokens under keys beginning with tokenizer.. The inspector displays these values as ordinary metadata, including arrays. It does not tokenize text, verify vocabulary behavior, or use tokenizer entries in its RAM calculation.

Large arrays are collapsed in the interface to keep the report usable. Expanding an entry reveals the parsed values, and the JSON download retains them.

Displayed versus interpreted fields

Keys outside the calculation list are still useful for comparing conversions and recording provenance. Their meaning comes from the GGUF producer and the relevant model architecture, not from AI Model Inspector. The tool does not maintain an architecture-specific dictionary for every possible GGUF key.

Numeric-looking metadata is used by the estimator only when it is positive and finite. A field being visible does not guarantee that it was complete, current, or honored by a particular inference runtime.

Open the inspector to search metadata from your file. See inference memory for the formula that consumes the fields above.