Blog · October 5, 2026 · by Maxime Carriere

What is GGUF? How on-device language models routed around the standard

Run an open language model on your own machine, rather than behind an API, and the odds are good it arrives as a GGUF file. Ollama loads it, LM Studio loads it, and so does the X-ray box: its report writer, Google’s MedGemma, is one 2.49 GB file beside the vision model on a $249 Jetson. We opened that file byte by byte for this post, and on the way found a library default that made the box three times slower without saying a word.

GGUF is a binary model format designed to be loaded directly by inference engines, in practice mostly for language models. It holds the weights, plus what a program needs to use them: the name of the architecture, the tokenizer (which turns text into token IDs, the numbers the model reads) and the chat template (the exact wrapping a conversation needs before the model sees it). It comes from llama.cpp, an open source engine for running language models, which is also its reference implementation.

How it differs from ONNX

The previous post covered ONNX, the format most vision models pass through on their way to an edge chip. The two get listed side by side as model formats, but they store different things for different readers.

ONNX GGUF
What the file stores a graph of operations, and the weights the weights, and metadata. No graph
Where the architecture is implemented represented as a graph in the file implemented by the runtime; the file names which architecture to instantiate
How compressed (quantised) weights are stored several ways: extra graph operations, 4 bit types (since 2024), or the vendor compiler’s job as a storage type on each weight array, packed in fixed blocks
Who defines it a specification under the Linux Foundation a specification by the ggml and llama.cpp maintainers, no standards body
Who reads it vendor compilers, ONNX Runtime llama.cpp and engines built on its library, plus a few others
Typical models vision models: CNNs, detectors, vision transformers language models, and vision-language models with their image encoder in a second file

The key difference is the second row. A graph is the architecture written out: every mathematical operation the input passes through, in order. An ONNX file contains it, so a compiler that has never seen the model can still build it. A GGUF file contains only the architecture’s name, gemma3, and llama.cpp builds the graph from its own C++ code for that name.

So llama.cpp has to implement every architecture it runs: about 150, by the architecture list in its source tree, October 2026. A fine-tuned or custom-trained model loads as long as its structure is one of those: new weights in a known layout. A genuinely new architecture does not load until someone writes its code into llama.cpp, along with the converter that produces its GGUF.

llama.cpp did not make ONNX its interchange format for language models. It kept the architecture in the runtime and defined its own file format for the weights and metadata, and most on-device language models followed it.

Two rows, file on the left and the program that reads it on the right. ONNX: the file contains a small chain of operations, Conv, Relu and so on, plus 32 bit weights, and is read by a vendor compiler that needs no prior knowledge of the model. GGUF: the file contains only the label architecture equals gemma3, a tokenizer and chat template, and a shorter strip of roughly 5 bit weights; it is read by llama.cpp, which holds the chain of operations itself, one per architecture, written in C++.
Both get a trained model onto a device. In ONNX the architecture travels with the file. In GGUF it stays in the runtime, and the file only names it.

Four names, one pipeline

Four boxes in a row joined by arrows. Llama, Gemma, Qwen: the model, trained weights, from Meta, Google and Alibaba. Converted to GGUF: the file format, specified by the ggml and llama.cpp maintainers. Loaded by llama.cpp: the engine, an open source project. Used in Ollama and LM Studio: apps that run it, on llama.cpp or its ggml library.
Llama is a model, GGUF a file, llama.cpp an engine, Ollama an app. Only one of them is a file.

Llama is Meta’s model family, and llama.cpp began as a plain C++ implementation of it before growing to run Gemma, Qwen and most other open architectures. Ollama started as an app around llama.cpp, and now runs many models on its own engine, built on ggml, the library underneath llama.cpp.

What is inside the file

Conceptually, a GGUF file has four parts, laid out in order.

A horizontal strip representing the 2.49 GB file to scale. The header, metadata and tensor table together are a sliver of 6.5 MB at the start, 0.3 percent. Then the weights: 1.53 GB of Q4_K tensors, 0.40 GB of Q6_K tensors, and 0.55 GB for the token embedding, also Q6_K. Below, a zoom on the sliver lists metadata keys with plain English beside each: architecture gemma3, which architecture to build; size label 4B; 262,208 tokens, the vocabulary; a 1,532 character chat template, the conversation format; context length 131,072; and the calibration text the quantiser used.
Almost all of it is weights: everything else is the sliver at the far left. The file the X-ray box loads, read by a short script with no dependencies; its hash matches the published file byte for byte.
  • The header: what kind of file this is. Four magic bytes, a format version, and how many entries follow.
  • The metadata: what a program needs to serve the model. Key and value pairs: the architecture, the context length, the full 262,208 token vocabulary, the chat template.
  • The tensor table: a directory of the weights. One entry per tensor, a named array of weights, giving its shape, its storage type and where its bytes sit.
  • The data: the weights themselves. Padded to fixed boundaries, so a program can map the file into memory and use the bytes where they lie.

Why language models needed their own format

The format addresses three problems that grow with model size.

  • Quantisation is a storage type. Quantised weights are stored with fewer bits per number, smaller and slightly less exact. Every tensor records how its numbers are packed, so the file people download is already the compact file the inference runtime can consume directly.
  • The weights can be memory-mapped without copying. The program points at the file instead of copying it, and the operating system fetches each part from disk when first needed. Where a backend supports it, the inference kernels work on the packed representation directly. A model whose file is larger than available RAM can still be mapped, but every token passes through every weight, so the file is reread from disk for each one: seconds per token.
  • Most of what the runtime needs travels in one file. The weights, the tokenizer, the chat template and the architecture name, with no folder of configuration files beside it. What does not travel is the architecture’s implementation, which lives in llama.cpp. Two further exceptions: a vision-language model ships its image encoder as a second file, and very large models are split into numbered parts.
Four horizontal bars for the same model at different quantisations, with bits per weight on the right: BF16, 7.77 GB, 16.0 bits; Q8_0, 4.13 GB, 8.5 bits; Q4_K_M, 2.49 GB, 5.1 bits, highlighted as the one on the box; Q4_K_S, 2.38 GB, 4.9 bits. Below, one Q4_K block drawn to scale: 2 bytes scaling the scales, 2 bytes scaling the minimums, 12 bytes holding 8 scales and 8 minimums at 6 bits each, then 128 bytes holding 256 weights at 4 bits each: 144 bytes for 256 weights, 128 bytes of 4 bit codes plus 16 bytes of scales and minimums, or 4.5 bits per weight.
One model, four published files; BF16 is the unquantised 16 bit original. Bits per weight are file size over parameter count, sizes as published on Hugging Face, October 2026.

A type such as Q4_K is a byte layout. Weights are stored in blocks of 256, each weight as a 4 bit code, and every block carries the scales and minimums that turn codes back into numbers. A block is 144 bytes for 256 weights: 128 bytes of codes plus 16 bytes of scales and minimums, or 4.5 bits per weight.

The name on a file is a recipe that mixes types, M for a medium mix, S for a small one. Our Q4_K_M file stores 204 tensors in Q4_K and raises 35 to the larger Q6_K, picked by a fixed rule as the ones usually most sensitive to rounding. One is the token embedding, the lookup table that maps each token ID to a vector of numbers, which alone is 17 percent of the model.

So this four bit model spends 5.1 bits per weight on average. The BF16 file, the 16 bit original, is three times larger and would not fit beside the vision model on an 8 GB board.

The rounding can be steered. Before quantising, a separate tool runs sample text through the model and records how strongly each column of each weight matrix is used on that text, a proxy for which weights are most sensitive to rounding, and the quantiser rounds those more carefully. The record is called an importance matrix. This file’s came from unsloth_calibration_medgemma-4b-it.txt, the converter’s text rather than radiology reports. Whether that matters for a product is a measurement, not an assumption.

One project in charge: fast to change, limited in reach

Because the same project defines the format and runs it, the format changes as fast as the code. A new quantisation type is usable as soon as it is merged into llama.cpp: there is no standards committee to approve it, and no other company’s software has to add support first. ONNX moves more slowly for exactly that reason.

A table of chips in three groups. Runs today: any x86 or Arm CPU; NVIDIA GPUs including Jetson, through CUDA; Apple M-series, through Metal; AMD GPUs, through HIP; Intel GPUs, through SYCL; other GPUs including Qualcomm's Adreno, through Vulkan or OpenCL; Huawei Ascend NPUs, through CANN. Early, inside llama.cpp: Intel NPUs through OpenVINO; Snapdragon NPUs through Hexagon, experimental. No NPU path: the Rockchip NPU, TI's C7x accelerator, NXP's i.MX NPU and Hailo-8, each reachable only through its vendor's own toolchain, so llama.cpp falls back to a CPU path.
From llama.cpp's backend documentation, October 2026, main backends only. The bottom group is the NPUs of the edge vendors from the ONNX post.

The cost is coverage. A GGUF file runs only where llama.cpp has a backend, its own code for that type of chip, and no vendor toolchain from the ONNX post reads the format.

llama.cpp has no backend for the neural processing units (NPUs) on Rockchip, TI, NXP or Hailo hardware, checked against its source tree in October 2026. Without one, it falls back to a CPU path, and the accelerator the board was chosen for sits idle. Rockchip’s own route to its NPU, RKLLM, converts models from Hugging Face’s format into its own, not from GGUF.

NPU support is arriving from one side only. Intel’s NPUs, through OpenVINO, and Qualcomm’s, through Hexagon, are now reachable by backends inside llama.cpp, not by GGUF importers in the vendors’ toolchains. Both are early. The Hexagon backend is marked experimental and accepts four storage types, Q4_0, Q8_0, MXFP4 and F32, so most of a Q4_K_M file like ours would run on the CPU.

Pitfalls, and the alternatives

Three things go wrong in practice:

  • The runtime is older than the file. A llama.cpp build without code for the file’s architecture refuses it with unknown model architecture. Python bindings bundle their own copy of llama.cpp, so the version of a pip package decides which models load.
  • The file is somebody else’s snapshot. Most GGUF files are converted by third parties, Unsloth in our case, who fix the template, tokenizer and quantisation on the day they convert. Corrected files can be re-uploaded under the same name, so pin the hash.
  • The GPU sits idle. In llama-cpp-python, the Python wrapper the box uses, the default n_gpu_layers=0 means none of the 34 layers is moved, or offloaded, to the GPU. No error, the same answer, and on the box 11.2 seconds instead of 3.3.

The alternatives split by who owns the hardware path, as of October 2026. Safetensors, Hugging Face’s format, is where the weights start, and nearly every GGUF is converted from one. MLX is Apple’s route on Apple silicon. ExecuTorch, from Meta, reaches the NPUs in Qualcomm, MediaTek, Apple and NXP chips that GGUF mostly cannot. ONNX Runtime GenAI, from Microsoft, puts language models back through ONNX. TensorRT-LLM is NVIDIA’s own stack, for NVIDIA GPUs only.

Each of those has a large company or foundation behind it. GGUF has a repository and its maintainers: it can gain a quantisation type in a week, and drop one without asking anyone, as it did with several ARM-specific types in 2024. How much that matters depends on how long the product has to be supported.

Where we fit

The two formats divide the work we do. Vision models on an accelerator arrive as ONNX and leave as a vendor binary. Language models on the same box arrive as GGUF and run on llama.cpp, and the X-ray reader does both at once, with nothing leaving the room.

Most of the failures above are a default taken by somebody else: a converter’s recipe and calibration text, a library’s offload setting.

Three boxes in a row joined by arrows: the published weights, GGUF, and llama.cpp on the box, with a Kernwerk badge on each arrow. Under the first arrow, making the file: pick the layout the memory budget allows, calibrate on text from your domain, measure the accuracy on the task. Under the second arrow, loading the file: pin the file by its hash, check the template it applies, check every layer is offloaded.
The same two arrows as the ONNX post. Making the file sets the best accuracy you can get. Loading it decides whether you get it, and how fast.

Running a language model on hardware you ship, rather than on somebody’s API? We do the work between the file and the chip. Talk to us.