GGUF reference
What AI Model Inspector reads from a GGUF file
GGUF is a binary container used to store model information and tensor data. The inspector reads the descriptive part of the file in your browser; it does not load or run the model.
How a GGUF file is arranged
The header starts with the GGUF signature, a format version, a tensor count, and a metadata-entry count. Metadata follows as typed key-value pairs. Next comes a descriptor for each tensor. The large tensor-data region begins at an aligned byte offset after those descriptors.
A tensor descriptor contains a name, a list of dimensions, a numeric GGML tensor type, and an offset relative to the tensor-data region. The dimensions describe the tensor shape. For example, two dimensions might represent the rows and columns of a weight matrix.
What the inspector extracts
AI Model Inspector accepts GGUF versions 2 and 3. It reads all metadata scalar types defined in the parser: signed and unsigned 8-, 16-, 32-, and 64-bit integers, 32- and 64-bit floats, booleans, UTF-8 strings, and arrays of supported values. Large integer values are shown as decimal strings so the browser does not round them.
The report shows header counts, the byte position where tensor data starts, every metadata entry, and tensor descriptors. It reads the file in 4 MiB chunks while parsing the descriptive section, so it does not copy the complete tensor payload into memory for GGUF inspection.
What inspection can tell you
Metadata can identify the model architecture, declared context length, embedding size, block count, attention layout, tokenizer, and other conversion details when the file supplies those keys. Tensor descriptors show which named tensors are present, their shapes, storage types, and offsets. Together, these details can help you check a conversion, compare two variants, and understand the inputs to the memory estimate.
The tool also counts parameters from tensor shapes when every dimension can be handled safely. That count is used for alternative quantization projections; it is not taken from a model name or guessed parameter label.
Current limits
The parser validates the signature and rejects versions other than 2 and 3. It does not decode tensor payloads, validate weights, execute inference, or translate numeric GGML tensor-type codes into names. A valid header and descriptor table do not prove that every payload byte is intact.
Counts and lengths larger than JavaScript's safe integer range are rejected. Invalid UTF-8 metadata strings, unsupported metadata value types, truncated files, and invalid boolean encodings also stop parsing. Inspection happens in the main page thread, so unusually large or malformed descriptive sections can still be slow.
To inspect a file, open the tool and choose a .gguf file. Read GGUF metadata for the keys used by the memory estimator.