What is GGUF? How on-device language models routed around the standard
Run an open language model on your own machine, rather than behind an API, and the odds are good it arrives as a GGUF file. Ollama loads it, LM Studio loads it, and so does the X-ray box: its report writer, Google’s MedGemma, is one 2.49 GB file beside the vision model on a $249 Jetson. We opened that file byte by byte for this post, and on the way found a library default that made the box three times slower without saying a word.
GGUF is a binary model format designed to be loaded directly by inference engines, in practice mostly for language models. It holds the weights, plus what a program needs to use them: the name of the architecture, the tokenizer (which turns text into token IDs, the numbers the model reads) and the chat template (the exact wrapping a conversation needs before the model sees it). It comes from llama.cpp, an open source engine for running language models, which is also its reference implementation.
How it differs from ONNX
The previous post covered ONNX, the format most vision models pass through on their way to an edge chip. The two get listed side by side as model formats, but they store different things for different readers.
| ONNX | GGUF | |
|---|---|---|
| What the file stores | a graph of operations, and the weights | the weights, and metadata. No graph |
| Where the architecture is implemented | represented as a graph in the file | implemented by the runtime; the file names which architecture to instantiate |
| How compressed (quantised) weights are stored | several ways: extra graph operations, 4 bit types (since 2024), or the vendor compiler’s job | as a storage type on each weight array, packed in fixed blocks |
| Who defines it | a specification under the Linux Foundation | a specification by the ggml and llama.cpp maintainers, no standards body |
| Who reads it | vendor compilers, ONNX Runtime | llama.cpp and engines built on its library, plus a few others |
| Typical models | vision models: CNNs, detectors, vision transformers | language models, and vision-language models with their image encoder in a second file |
The key difference is the second row. A graph is the architecture written out: every mathematical operation the input passes through, in order. An ONNX file contains it, so a compiler that has never seen the model can still build it. A GGUF file contains only the architecture’s name, gemma3, and llama.cpp builds the graph from its own C++ code for that name.
So llama.cpp has to implement every architecture it runs: about 150, by the architecture list in its source tree, October 2026. A fine-tuned or custom-trained model loads as long as its structure is one of those: new weights in a known layout. A genuinely new architecture does not load until someone writes its code into llama.cpp, along with the converter that produces its GGUF.
llama.cpp did not make ONNX its interchange format for language models. It kept the architecture in the runtime and defined its own file format for the weights and metadata, and most on-device language models followed it.
Four names, one pipeline
Llama is Meta’s model family, and llama.cpp began as a plain C++ implementation of it before growing to run Gemma, Qwen and most other open architectures. Ollama started as an app around llama.cpp, and now runs many models on its own engine, built on ggml, the library underneath llama.cpp.
What is inside the file
Conceptually, a GGUF file has four parts, laid out in order.
- The header: what kind of file this is. Four magic bytes, a format version, and how many entries follow.
- The metadata: what a program needs to serve the model. Key and value pairs: the architecture, the context length, the full 262,208 token vocabulary, the chat template.
- The tensor table: a directory of the weights. One entry per tensor, a named array of weights, giving its shape, its storage type and where its bytes sit.
- The data: the weights themselves. Padded to fixed boundaries, so a program can map the file into memory and use the bytes where they lie.
Why language models needed their own format
The format addresses three problems that grow with model size.
- Quantisation is a storage type. Quantised weights are stored with fewer bits per number, smaller and slightly less exact. Every tensor records how its numbers are packed, so the file people download is already the compact file the inference runtime can consume directly.
- The weights can be memory-mapped without copying. The program points at the file instead of copying it, and the operating system fetches each part from disk when first needed. Where a backend supports it, the inference kernels work on the packed representation directly. A model whose file is larger than available RAM can still be mapped, but every token passes through every weight, so the file is reread from disk for each one: seconds per token.
- Most of what the runtime needs travels in one file. The weights, the tokenizer, the chat template and the architecture name, with no folder of configuration files beside it. What does not travel is the architecture’s implementation, which lives in llama.cpp. Two further exceptions: a vision-language model ships its image encoder as a second file, and very large models are split into numbered parts.
A type such as Q4_K is a byte layout. Weights are stored in blocks of 256, each weight as a 4 bit code, and every block carries the scales and minimums that turn codes back into numbers. A block is 144 bytes for 256 weights: 128 bytes of codes plus 16 bytes of scales and minimums, or 4.5 bits per weight.
The name on a file is a recipe that mixes types, M for a medium mix, S for a small one. Our Q4_K_M file stores 204 tensors in Q4_K and raises 35 to the larger Q6_K, picked by a fixed rule as the ones usually most sensitive to rounding. One is the token embedding, the lookup table that maps each token ID to a vector of numbers, which alone is 17 percent of the model.
So this four bit model spends 5.1 bits per weight on average. The BF16 file, the 16 bit original, is three times larger and would not fit beside the vision model on an 8 GB board.
The rounding can be steered. Before quantising, a separate tool runs sample text through the model and records how strongly each column of each weight matrix is used on that text, a proxy for which weights are most sensitive to rounding, and the quantiser rounds those more carefully. The record is called an importance matrix. This file’s came from unsloth_calibration_medgemma-4b-it.txt, the converter’s text rather than radiology reports. Whether that matters for a product is a measurement, not an assumption.
One project in charge: fast to change, limited in reach
Because the same project defines the format and runs it, the format changes as fast as the code. A new quantisation type is usable as soon as it is merged into llama.cpp: there is no standards committee to approve it, and no other company’s software has to add support first. ONNX moves more slowly for exactly that reason.
The cost is coverage. A GGUF file runs only where llama.cpp has a backend, its own code for that type of chip, and no vendor toolchain from the ONNX post reads the format.
llama.cpp has no backend for the neural processing units (NPUs) on Rockchip, TI, NXP or Hailo hardware, checked against its source tree in October 2026. Without one, it falls back to a CPU path, and the accelerator the board was chosen for sits idle. Rockchip’s own route to its NPU, RKLLM, converts models from Hugging Face’s format into its own, not from GGUF.
NPU support is arriving from one side only. Intel’s NPUs, through OpenVINO, and Qualcomm’s, through Hexagon, are now reachable by backends inside llama.cpp, not by GGUF importers in the vendors’ toolchains. Both are early. The Hexagon backend is marked experimental and accepts four storage types, Q4_0, Q8_0, MXFP4 and F32, so most of a Q4_K_M file like ours would run on the CPU.
Pitfalls, and the alternatives
Three things go wrong in practice:
- The runtime is older than the file. A llama.cpp build without code for the file’s architecture refuses it with
unknown model architecture. Python bindings bundle their own copy of llama.cpp, so the version of a pip package decides which models load. - The file is somebody else’s snapshot. Most GGUF files are converted by third parties, Unsloth in our case, who fix the template, tokenizer and quantisation on the day they convert. Corrected files can be re-uploaded under the same name, so pin the hash.
- The GPU sits idle. In llama-cpp-python, the Python wrapper the box uses, the default
n_gpu_layers=0means none of the 34 layers is moved, or offloaded, to the GPU. No error, the same answer, and on the box 11.2 seconds instead of 3.3.
The alternatives split by who owns the hardware path, as of October 2026. Safetensors, Hugging Face’s format, is where the weights start, and nearly every GGUF is converted from one. MLX is Apple’s route on Apple silicon. ExecuTorch, from Meta, reaches the NPUs in Qualcomm, MediaTek, Apple and NXP chips that GGUF mostly cannot. ONNX Runtime GenAI, from Microsoft, puts language models back through ONNX. TensorRT-LLM is NVIDIA’s own stack, for NVIDIA GPUs only.
Each of those has a large company or foundation behind it. GGUF has a repository and its maintainers: it can gain a quantisation type in a week, and drop one without asking anyone, as it did with several ARM-specific types in 2024. How much that matters depends on how long the product has to be supported.
Where we fit
The two formats divide the work we do. Vision models on an accelerator arrive as ONNX and leave as a vendor binary. Language models on the same box arrive as GGUF and run on llama.cpp, and the X-ray reader does both at once, with nothing leaving the room.
Most of the failures above are a default taken by somebody else: a converter’s recipe and calibration text, a library’s offload setting.
Running a language model on hardware you ship, rather than on somebody’s API? We do the work between the file and the chip. Talk to us.