Blog · October 13, 2026 · by Maxime Carriere

Same weights, different answers: why you test a model on its board and certify what runs

A model file is not a behaviour. The earlier posts on ONNX and GGUF followed a model’s file to the chip; this one asks what the chip then computes. The same weights, run on two different processors, compute slightly different numbers, and sometimes those numbers lead to a different answer. The cause is that floating-point addition is not associative: each processor’s kernels accumulate the model’s dot products in a different order, so the rounding errors, and the results, differ in the last bits.

The experiment

One model, two devices, and as little else as possible allowed to vary.

  • Model: Qwen2.5-1.5B-Instruct, the same file on both machines, checked by its SHA-256 hash, a fingerprint that changes if a single bit does.
  • Devices: a MacBook with an Apple M1, running PyTorch on its graphics processor (GPU), and an NVIDIA Jetson Orin Nano Super, running PyTorch on CUDA, NVIDIA’s GPU platform.
  • Conditions: FP, 16-bit floating point (fp16) weights and computation on both devices, with the same model file and input token IDs; the hardware, backend and kernels differ. INT8, the Jetson runs TensorRT-LLM with 8-bit weights and 16-bit activations.
  • Benchmarks: fixed subsets, not random samples, chosen to fit the Jetson’s run time. MMLU (Massive Multitask Language Understanding): the first 20 questions of each of its 57 subjects, 1,140 in all. HellaSwag: the first 1,000 validation items. GSM8K (Grade School Math): the first 250 of 1,319 test problems, answered in writing. Prompts and scoring are those of lm-eval, EleutherAI’s evaluation harness, 0-shot for the first two and 5-shot for GSM8K.
  • Controls: prompts tokenized once and fed to both devices as the same token IDs; greedy decoding, which always takes the highest-scoring next token, and batch size 1. Each setup was rerun, all MMLU and HellaSwag items and 50 GSM8K problems, with bit-identical results, so the differences below are not run-to-run noise.

Similar scores, different answers

Three panels comparing a MacBook M1 and a Jetson Orin Nano Super running Qwen2.5-1.5B with the same weights and input tokens, in two conditions: fp16 on both devices, and int8 weights through TensorRT-LLM on the Jetson. Panel A, vertical bars, share of questions answered differently, with the score change below: MMLU 0.09 percent and 9.8 percent, plus 0.1 and minus 0.3 percentage point; HellaSwag 0.2 and 2.3 percent, plus 0.1 and minus 0.1 point; GSM8K 7.2 and 53.2 percent, minus 0.4 and minus 11.2 points. Panel B, vertical bars on a log scale, mean gap between the devices: first-token logits 0.038 and 1.86; MMLU answer-choice log-likelihoods 0.008 and 0.29; HellaSwag 0.019 and 0.62. Panel C, two scatter plots of every one of the 4,560 MMLU answer choices, MacBook log-likelihood against Jetson log-likelihood. With fp16 the points lie on the diagonal, largest gap 0.06, and one answer changes. With int8 they spread into a cloud around the diagonal, largest gap 7.5, and the choices of the 112 questions whose answer changed, drawn in red, cluster where scores are close to zero, among the model's leading choices.
Both conditions are compared with the MacBook in fp16. A: share of questions whose extracted answer differs, and the change in score in percentage points. B: mean absolute gap in raw logits or log-likelihoods, in nats; for the first token, the largest gap over the full vocabulary per prompt, averaged over 200 MMLU prompts; for MMLU and HellaSwag, the gap of each answer choice's log-likelihood, averaged over all choices. C: every MMLU answer choice, one dot each.

Three measures are easy to confuse, so we keep them apart. Answer disagreement is the share of questions whose extracted answer differs between the devices. Accuracy change is the difference in the share judged correct. Text disagreement counts generated texts that differ at all, even with the same final answer.

Panel B shows why answers move at all: every number the model computes shifts slightly between the devices, and on these measurements the int8 path increases the mean gaps by roughly 30 to 50 times, depending on the metric. Panel C shows it for every MMLU answer choice: with fp16 the 4,560 log-likelihoods sit on the diagonal, the largest gap 0.06; with int8 they spread, by up to 7.5.

Panel A shows what that does to answers. With fp16 on both devices, the scores agree to within half a percentage point. That is the number most evaluations stop at, and it hides what happened underneath:

  • Multiple choice barely moves. One MMLU answer in 1,140 differs, and two HellaSwag answers in 1,000.
  • Written reasoning moves more. On GSM8K the generated text differs on 38 of 250 problems, and the final answer on 18 of them, 7.2%. An answer runs to about 120 tokens, which gives a small difference many chances to change a word, and one changed word sends the rest of the solution down another path.
  • The int8 path moves most. 9.8% of MMLU answers change, yet the score moves by 0.3 points: of the 112 that changed, 40 became wrong, 37 became right and the rest went from one wrong answer to another. On GSM8K, half the answers change and the score drops 11 points. Running fp16 and int8 on the same Jetson gives the same picture, so this comes from the int8 path rather than the chip. That path changes two things at once, the weights’ precision and the TensorRT runtime, and without a TensorRT fp16 run we cannot say how much each contributes.

We have not isolated the cause of the 11-point drop. The int8 answers stay fluent but make arithmetic slips, and break the expected answer format about twice as often as fp16, on 60 of the 250 problems against 31. Whatever the cause, it is what this toolchain produced on this board, which is the argument for measuring there.

Why the hardware changes the numbers

Floating-point addition is not associative: (a + b) + c and a + (b + c) can round to different results. Every layer of a model sums thousands of products, and each processor’s kernels order those sums differently, fuse different operations and accumulate at different precisions.

The result is a tiny, deterministic difference. None of the 200 output vectors we compared, the logits of the first 200 MMLU prompts, was bit-identical across the devices: the average entry differed by 0.006 and the largest by 0.07, on logits whose largest values are in the tens.

A difference that small matters only where the leading scores are close. At each step greedy decoding takes the token with the highest logit, the raw score the model gives each candidate. If the top two logits are equal, or one fp16 rounding step apart, a different rounding on the other device is enough to put the other token first.

One tie, two answers: GSM8K maths problem 249, the same model in fp16 on both devices, in four steps. Step 1, both devices get the same question: Sue ate 4 times as many cookies as her sister on Monday and twice as many on Tuesday; her sister ate 5 on Monday and 13 on Tuesday; a cookie has 200 calories; how many more calories did Sue eat than her sister? The correct answer is 5,600. Step 2, both write the same first sentence: On Monday, Sue ate 4 times 5 equals 20 cookies. Step 3, the next word is a tie: on the MacBook the candidates On and Her have equal scores. A tie is broken by the last bits of the arithmetic; the MacBook picks On, and on the Jetson rounding puts Her one step ahead. Step 4, each device continues from its own word. The MacBook M1 writes On Tuesday, Sue ate 2 times 13 equals 26 cookies, then Sue 46, sister 18, 46 minus 18 times 200 equals 5,600: correct. The Jetson Orin Nano writes Her sister ate 5 plus 13 equals 18 cookies, then Sue 38, sister 36, 2 times 200 equals 400: wrong. Below, how rare a tie is: all 30,388 words generated on the MacBook sorted by how close their top two candidates scored, 38 exact ties, 73 within 1/64, then 226, 836, 2,978, 7,286 and 18,951 in wider bins. Only 111, 0.4 percent, were ties or within one rounding step of a tie, and all 38 splits coincided with one of them in the MacBook trace.
From the study's own outputs. The first 23 tokens are identical on both devices; one tie later, they answer 5,600 and 400.

The histogram supports that mechanism. Of 30,388 generated steps, only 111 were within one fp16 rounding step of a tie, 0.4%. All 38 divergences coincided with one of those steps in the MacBook’s trace: 23 exact ties, 15 one step apart.

We did not log the Jetson’s logits, so this is evidence for the mechanism rather than a full causal reconstruction. In this run, no split happened where the leading scores were far apart: the hardware mattered where the scores were close, and one changed token sent the rest of the solution down another path.

Two consequences follow. Aggregate scores hide this, because close calls happen mostly on questions the model already gets wrong: of the 18 GSM8K answers that changed in fp16, 17 went from one wrong answer to another. And a benchmark run on a laptop, or on a cloud GPU, does not certify what the same file will answer on the board.

From accident to attack

So far the differences are an accident. Two recent papers show they can be engineered. Three levels of evidence need keeping apart: ordinary divergence between platforms, which our study measures; attacks that exploit it, which the papers demonstrate on their own models and platforms; and the risk to a particular edge deployment, which neither establishes.

  • Inputs that flip on one backend. Möller et al. (International Conference on Machine Learning, ICML 2025) construct “Chimera examples”: inputs that one image classifier labels differently depending on the linear algebra library underneath, such as Intel’s Math Kernel Library, Apple Accelerate or NVIDIA cuBLAS. They found them for every pair of the six libraries tested, with search success rates from about 9% to 100%, highest whenever cuBLAS was one of the pair. What it buys an attacker is an input that behaves as expected on the system where a model is developed and tested, and differently on the production or embedded system where it runs. The attack assumes the attacker knows the model; the authors’ keyed-noise defence cuts success to under 1%.
  • A model that misbehaves on one platform. Loose et al. (FloatDoor, arXiv preprint, June 2026) fine-tune a model with two small adapters, known as LoRA, so that it amplifies platform differences into a signal it can read, then condition its behaviour on that signal. No special input is needed. What it buys an attacker is a model that passes its audit on one platform and misbehaves on another: it can silently mark which platform it runs on, or write insecure code only there. In their case study, one Qwen3-8B checkpoint writes exploitable code in 49% of the programs it produces for a set of coding prompts on an NVIDIA A100, and in 16% on an NVIDIA H200, while general benchmarks move by under 2 points. Because platforms map to organisations and regions, the authors note, the same trick could target one group of users with, for example, biased summaries.
One checkpoint, Qwen3-8B with two fine-tuning adapters merged in, one set of weights and one hash, is run on two GPUs with the same coding prompts, from Loose et al., arXiv preprint, June 2026. On the auditor platform, an NVIDIA H200 where the model is checked, 15.7 percent of programs, plus or minus 3.3, contain an exploitable flaw. On the target platform, an NVIDIA A100 where the model is served, 49.0 percent, plus or minus 9.0. A dashed line marks 11.8 percent, the same model before the second adapter.
FloatDoor, table 4. The audit passes where the audit is run. What matters is what the model does where it is served.

The authors name the weakness precisely: a time-of-check, time-of-use gap between where a model is audited and where it is served.

On their fingerprinting model, running in 32-bit floating point or pruning 10% of the weights removed the platform signal at under one point of benchmark cost. But 32-bit inference doubles the memory and slows generation, which an 8 GB board can rarely afford, and the authors leave an attacker who anticipates these defences to future work.

Until such defences are standard, they conclude, trusted model supply chains are the only viable safeguard. We would put it more narrowly: provenance is necessary but not sufficient, since a model from an authenticated source can still be malicious, and it works alongside testing on the serving platform and checking what runs.

FloatDoor’s backdoor was shown on datacentre GPUs, Google tensor processing units (TPUs) and server CPUs. It has not been demonstrated on an edge board. Its divergence measurements cover 23 platforms, including the Apple M1 GPU used here, and the only pairs it could not tell apart were server CPUs falling back to 32-bit arithmetic.

An edge deployment usually has two platforms that round differently: the evaluation machine and the board, two board generations, or, by the same mechanism, one board before and after a software upgrade. Earlier work hides backdoors that fire only after quantization (Hong et al., Conference on Neural Information Processing Systems, NeurIPS 2021) or after compilation (Chen et al., 2025 preprint), exactly the steps an edge deployment adds.

What this means for an edge deployment

The model you evaluated and the model your device runs may compute different answers, even when they share the same weights. Three checks close the gaps.

Each check covers a failure the others miss. FloatDoor passes any identity check by design, since the audited and the served model share one hash. What it defeats is testing on the wrong platform, and testing on the right one catches it only if your test set covers the behaviour the attacker chose. Proof of identity catches the opposite failure: a model, runtime or board swapped after the test. Neither says whether the weights were trustworthy to begin with, which is the authors’ point about supply chains.

Limits of this study

This is a small study: one model, one pair of devices, and fixed benchmark subsets rather than random samples, so the rates above describe these items and should not be read as estimates for the full benchmarks.

The 7.2% uses the flexible answer extraction of lm-eval, the evaluation harness we used; the strict one gives 5.6%. The int8 condition quantizes weights only. A deployment that also quantizes activations could behave differently; we have not measured that configuration. TensorRT’s own fp16 kernels could not be tested separately, because that engine did not fit in the Jetson’s 8 GB.

The tie histogram uses the MacBook’s logits; the Jetson’s were not logged. Versions: macOS 26.6 and PyTorch 2.11 on the MacBook; JetPack 6.2 in MAXN_SUPER power mode, PyTorch 2.11 and CUDA 12.6 on the Jetson, and TensorRT-LLM 0.12 on TensorRT 10.4 in NVIDIA’s container for the int8 runs.

The attack numbers are the papers’ own, on their models and platforms, as of October 2026. The prompts and token IDs, per-item outputs and disagreement labels, and the scripts behind every figure are in our study repository, and we share them on request.

Where we fit

Kernwerk does these checks, and delivers them as three things.

  • A behavioural report: what the model answers on the board you ship, item by item, against your reference, with every flipped answer listed.
  • Artifact identity: exactly which weights, engine, plugins, tokenizer and runtime were tested, by their hashes, after checking the weights against their publisher’s hash, and signed.
  • Device attestation: evidence that a deployed board loaded those measured artifacts.
What Kernwerk delivers, illustrative. Your model, with its publisher's hash checked, goes onto the board you ship. On the board, two checks. One, tested on the board: every answer compared with your reference, shown as a row of green squares for the same answer and one red square for a flipped answer, which is listed in the report. Two, certified what runs: the engine, plugins, tokenizer and runtime are hashed, signed, and attested by the board at load. Both feed one signed report stating which model by its hashes, which board and software, what it answered there, and what changed, item by item.
Illustrative. The test says what the model does on that board; the certificate says that board runs that model. The report binds the two.

How the attestation works

The attestation is an architecture we set up per deployment, not a feature every board ships with. Secure boot checks the firmware and operating system before they start. It says nothing about a model file loaded later.

So our loader measures each artifact as it loads it and extends its hash into a platform configuration register (PCR) of a firmware TPM. On a Jetson Orin that TPM runs inside OP-TEE, its trusted execution environment, once it has been enabled and provisioned; it is a configuration step, not a property of every board.

For a verifier to trust the result, three things must hold, and signing a list of hashes proves none of them:

  • The loader is trusted. It is itself measured by the boot chain before it runs, so a modified loader shows up in the registers.
  • The measurements are bound to the evidence. The TPM signs a quote over those register values and a fresh challenge from the verifier, and the verifier replays the measurement log against the quoted values.
  • The device identity is provisioned. The TPM’s attestation key is created and certified for that board, so the verifier knows which device produced the quote.

The quote then shows what was measured since boot on that device. It does not watch the model while it runs; that is runtime integrity monitoring, a separate problem.

On NXP and Rockchip parts the same chain needs OP-TEE hosting a firmware TPM, or a secure element, and how much of it a given board supports depends on its fusing and firmware. The artifact itself is signed as described before, and the TrustZone post covers the trusted execution side.

We do not sell silicon, so nothing stops the report from saying the board changed your model’s answers, and by how much.

Shipping a model to a device you do not control? We test it on the board it will run on, and make the board prove it runs that model. Talk to us.

References

  • Möller, J., Pirch, L., Weissberg, F., Baunsgaard, S., Eisenhofer, T. and Rieck, K. “Adversarial Inputs for Linear Algebra Backends.” Proceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267, 2025. Peer-reviewed.
  • Loose, N., Sander, J., Mächtle, F. and Eisenbarth, T. “FloatDoor: Platform-Triggered Backdoors in LLMs.” arXiv:2606.19535, June 2026. Preprint. arxiv.org/abs/2606.19535
  • Hong, S., Panaitescu-Liess, M.-A., Kaya, Y. and Dumitras, T. “Qu-anti-zation: Exploiting Quantization Artifacts for Achieving Adversarial Outcomes.” Advances in Neural Information Processing Systems 34 (NeurIPS), 2021. Peer-reviewed.
  • Chen, S., Peng, J., He, Y., Yang, J. and Ray, B. “Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers.” arXiv:2509.11173, 2025. Preprint.
  • Qwen Team. Qwen2.5-1.5B-Instruct, Hugging Face, commit 989aa79. Benchmarks: MMLU (Hendrycks et al., ICLR 2021), HellaSwag (Zellers et al., ACL 2019), GSM8K (Cobbe et al., 2021), run with EleutherAI’s lm-evaluation-harness 0.4.13.