Same weights, different answers: why you test a model on its board and certify what runs
A model file is not a behaviour. The earlier posts on ONNX and GGUF followed a model’s file to the chip; this one asks what the chip then computes. The same weights, run on two different processors, compute slightly different numbers, and sometimes those numbers lead to a different answer. The cause is that floating-point addition is not associative: each processor’s kernels accumulate the model’s dot products in a different order, so the rounding errors, and the results, differ in the last bits.
The experiment
One model, two devices, and as little else as possible allowed to vary.
- Model: Qwen2.5-1.5B-Instruct, the same file on both machines, checked by its SHA-256 hash, a fingerprint that changes if a single bit does.
- Devices: a MacBook with an Apple M1, running PyTorch on its graphics processor (GPU), and an NVIDIA Jetson Orin Nano Super, running PyTorch on CUDA, NVIDIA’s GPU platform.
- Conditions: FP, 16-bit floating point (fp16) weights and computation on both devices, with the same model file and input token IDs; the hardware, backend and kernels differ. INT8, the Jetson runs TensorRT-LLM with 8-bit weights and 16-bit activations.
- Benchmarks: fixed subsets, not random samples, chosen to fit the Jetson’s run time. MMLU (Massive Multitask Language Understanding): the first 20 questions of each of its 57 subjects, 1,140 in all. HellaSwag: the first 1,000 validation items. GSM8K (Grade School Math): the first 250 of 1,319 test problems, answered in writing. Prompts and scoring are those of lm-eval, EleutherAI’s evaluation harness, 0-shot for the first two and 5-shot for GSM8K.
- Controls: prompts tokenized once and fed to both devices as the same token IDs; greedy decoding, which always takes the highest-scoring next token, and batch size 1. Each setup was rerun, all MMLU and HellaSwag items and 50 GSM8K problems, with bit-identical results, so the differences below are not run-to-run noise.
Similar scores, different answers
Three measures are easy to confuse, so we keep them apart. Answer disagreement is the share of questions whose extracted answer differs between the devices. Accuracy change is the difference in the share judged correct. Text disagreement counts generated texts that differ at all, even with the same final answer.
Panel B shows why answers move at all: every number the model computes shifts slightly between the devices, and on these measurements the int8 path increases the mean gaps by roughly 30 to 50 times, depending on the metric. Panel C shows it for every MMLU answer choice: with fp16 the 4,560 log-likelihoods sit on the diagonal, the largest gap 0.06; with int8 they spread, by up to 7.5.
Panel A shows what that does to answers. With fp16 on both devices, the scores agree to within half a percentage point. That is the number most evaluations stop at, and it hides what happened underneath:
- Multiple choice barely moves. One MMLU answer in 1,140 differs, and two HellaSwag answers in 1,000.
- Written reasoning moves more. On GSM8K the generated text differs on 38 of 250 problems, and the final answer on 18 of them, 7.2%. An answer runs to about 120 tokens, which gives a small difference many chances to change a word, and one changed word sends the rest of the solution down another path.
- The int8 path moves most. 9.8% of MMLU answers change, yet the score moves by 0.3 points: of the 112 that changed, 40 became wrong, 37 became right and the rest went from one wrong answer to another. On GSM8K, half the answers change and the score drops 11 points. Running fp16 and int8 on the same Jetson gives the same picture, so this comes from the int8 path rather than the chip. That path changes two things at once, the weights’ precision and the TensorRT runtime, and without a TensorRT fp16 run we cannot say how much each contributes.
We have not isolated the cause of the 11-point drop. The int8 answers stay fluent but make arithmetic slips, and break the expected answer format about twice as often as fp16, on 60 of the 250 problems against 31. Whatever the cause, it is what this toolchain produced on this board, which is the argument for measuring there.
Why the hardware changes the numbers
Floating-point addition is not associative: (a + b) + c and a + (b + c) can round to different results. Every layer of a model sums thousands of products, and each processor’s kernels order those sums differently, fuse different operations and accumulate at different precisions.
The result is a tiny, deterministic difference. None of the 200 output vectors we compared, the logits of the first 200 MMLU prompts, was bit-identical across the devices: the average entry differed by 0.006 and the largest by 0.07, on logits whose largest values are in the tens.
A difference that small matters only where the leading scores are close. At each step greedy decoding takes the token with the highest logit, the raw score the model gives each candidate. If the top two logits are equal, or one fp16 rounding step apart, a different rounding on the other device is enough to put the other token first.
The histogram supports that mechanism. Of 30,388 generated steps, only 111 were within one fp16 rounding step of a tie, 0.4%. All 38 divergences coincided with one of those steps in the MacBook’s trace: 23 exact ties, 15 one step apart.
We did not log the Jetson’s logits, so this is evidence for the mechanism rather than a full causal reconstruction. In this run, no split happened where the leading scores were far apart: the hardware mattered where the scores were close, and one changed token sent the rest of the solution down another path.
Two consequences follow. Aggregate scores hide this, because close calls happen mostly on questions the model already gets wrong: of the 18 GSM8K answers that changed in fp16, 17 went from one wrong answer to another. And a benchmark run on a laptop, or on a cloud GPU, does not certify what the same file will answer on the board.
From accident to attack
So far the differences are an accident. Two recent papers show they can be engineered. Three levels of evidence need keeping apart: ordinary divergence between platforms, which our study measures; attacks that exploit it, which the papers demonstrate on their own models and platforms; and the risk to a particular edge deployment, which neither establishes.
- Inputs that flip on one backend. Möller et al. (International Conference on Machine Learning, ICML 2025) construct “Chimera examples”: inputs that one image classifier labels differently depending on the linear algebra library underneath, such as Intel’s Math Kernel Library, Apple Accelerate or NVIDIA cuBLAS. They found them for every pair of the six libraries tested, with search success rates from about 9% to 100%, highest whenever cuBLAS was one of the pair. What it buys an attacker is an input that behaves as expected on the system where a model is developed and tested, and differently on the production or embedded system where it runs. The attack assumes the attacker knows the model; the authors’ keyed-noise defence cuts success to under 1%.
- A model that misbehaves on one platform. Loose et al. (FloatDoor, arXiv preprint, June 2026) fine-tune a model with two small adapters, known as LoRA, so that it amplifies platform differences into a signal it can read, then condition its behaviour on that signal. No special input is needed. What it buys an attacker is a model that passes its audit on one platform and misbehaves on another: it can silently mark which platform it runs on, or write insecure code only there. In their case study, one Qwen3-8B checkpoint writes exploitable code in 49% of the programs it produces for a set of coding prompts on an NVIDIA A100, and in 16% on an NVIDIA H200, while general benchmarks move by under 2 points. Because platforms map to organisations and regions, the authors note, the same trick could target one group of users with, for example, biased summaries.
The authors name the weakness precisely: a time-of-check, time-of-use gap between where a model is audited and where it is served.
On their fingerprinting model, running in 32-bit floating point or pruning 10% of the weights removed the platform signal at under one point of benchmark cost. But 32-bit inference doubles the memory and slows generation, which an 8 GB board can rarely afford, and the authors leave an attacker who anticipates these defences to future work.
Until such defences are standard, they conclude, trusted model supply chains are the only viable safeguard. We would put it more narrowly: provenance is necessary but not sufficient, since a model from an authenticated source can still be malicious, and it works alongside testing on the serving platform and checking what runs.
FloatDoor’s backdoor was shown on datacentre GPUs, Google tensor processing units (TPUs) and server CPUs. It has not been demonstrated on an edge board. Its divergence measurements cover 23 platforms, including the Apple M1 GPU used here, and the only pairs it could not tell apart were server CPUs falling back to 32-bit arithmetic.
An edge deployment usually has two platforms that round differently: the evaluation machine and the board, two board generations, or, by the same mechanism, one board before and after a software upgrade. Earlier work hides backdoors that fire only after quantization (Hong et al., Conference on Neural Information Processing Systems, NeurIPS 2021) or after compilation (Chen et al., 2025 preprint), exactly the steps an edge deployment adds.
What this means for an edge deployment
The model you evaluated and the model your device runs may compute different answers, even when they share the same weights. Three checks close the gaps.
Each check covers a failure the others miss. FloatDoor passes any identity check by design, since the audited and the served model share one hash. What it defeats is testing on the wrong platform, and testing on the right one catches it only if your test set covers the behaviour the attacker chose. Proof of identity catches the opposite failure: a model, runtime or board swapped after the test. Neither says whether the weights were trustworthy to begin with, which is the authors’ point about supply chains.
Limits of this study
This is a small study: one model, one pair of devices, and fixed benchmark subsets rather than random samples, so the rates above describe these items and should not be read as estimates for the full benchmarks.
The 7.2% uses the flexible answer extraction of lm-eval, the evaluation harness we used; the strict one gives 5.6%. The int8 condition quantizes weights only. A deployment that also quantizes activations could behave differently; we have not measured that configuration. TensorRT’s own fp16 kernels could not be tested separately, because that engine did not fit in the Jetson’s 8 GB.
The tie histogram uses the MacBook’s logits; the Jetson’s were not logged. Versions: macOS 26.6 and PyTorch 2.11 on the MacBook; JetPack 6.2 in MAXN_SUPER power mode, PyTorch 2.11 and CUDA 12.6 on the Jetson, and TensorRT-LLM 0.12 on TensorRT 10.4 in NVIDIA’s container for the int8 runs.
The attack numbers are the papers’ own, on their models and platforms, as of October 2026. The prompts and token IDs, per-item outputs and disagreement labels, and the scripts behind every figure are in our study repository, and we share them on request.
Where we fit
Kernwerk does these checks, and delivers them as three things.
- A behavioural report: what the model answers on the board you ship, item by item, against your reference, with every flipped answer listed.
- Artifact identity: exactly which weights, engine, plugins, tokenizer and runtime were tested, by their hashes, after checking the weights against their publisher’s hash, and signed.
- Device attestation: evidence that a deployed board loaded those measured artifacts.
How the attestation works
The attestation is an architecture we set up per deployment, not a feature every board ships with. Secure boot checks the firmware and operating system before they start. It says nothing about a model file loaded later.
So our loader measures each artifact as it loads it and extends its hash into a platform configuration register (PCR) of a firmware TPM. On a Jetson Orin that TPM runs inside OP-TEE, its trusted execution environment, once it has been enabled and provisioned; it is a configuration step, not a property of every board.
For a verifier to trust the result, three things must hold, and signing a list of hashes proves none of them:
- The loader is trusted. It is itself measured by the boot chain before it runs, so a modified loader shows up in the registers.
- The measurements are bound to the evidence. The TPM signs a quote over those register values and a fresh challenge from the verifier, and the verifier replays the measurement log against the quoted values.
- The device identity is provisioned. The TPM’s attestation key is created and certified for that board, so the verifier knows which device produced the quote.
The quote then shows what was measured since boot on that device. It does not watch the model while it runs; that is runtime integrity monitoring, a separate problem.
On NXP and Rockchip parts the same chain needs OP-TEE hosting a firmware TPM, or a secure element, and how much of it a given board supports depends on its fusing and firmware. The artifact itself is signed as described before, and the TrustZone post covers the trusted execution side.
We do not sell silicon, so nothing stops the report from saying the board changed your model’s answers, and by how much.
Shipping a model to a device you do not control? We test it on the board it will run on, and make the board prove it runs that model. Talk to us.
References
- Möller, J., Pirch, L., Weissberg, F., Baunsgaard, S., Eisenhofer, T. and Rieck, K. “Adversarial Inputs for Linear Algebra Backends.” Proceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267, 2025. Peer-reviewed.
- Loose, N., Sander, J., Mächtle, F. and Eisenbarth, T. “FloatDoor: Platform-Triggered Backdoors in LLMs.” arXiv:2606.19535, June 2026. Preprint. arxiv.org/abs/2606.19535
- Hong, S., Panaitescu-Liess, M.-A., Kaya, Y. and Dumitras, T. “Qu-anti-zation: Exploiting Quantization Artifacts for Achieving Adversarial Outcomes.” Advances in Neural Information Processing Systems 34 (NeurIPS), 2021. Peer-reviewed.
- Chen, S., Peng, J., He, Y., Yang, J. and Ray, B. “Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers.” arXiv:2509.11173, 2025. Preprint.
- Qwen Team. Qwen2.5-1.5B-Instruct, Hugging Face, commit 989aa79. Benchmarks: MMLU (Hendrycks et al., ICLR 2021), HellaSwag (Zellers et al., ACL 2019), GSM8K (Cobbe et al., 2021), run with EleutherAI’s lm-evaluation-harness 0.4.13.