What is actually running inside your earbuds?
You talk to an AI through your earbuds every day. Almost none of the large-model intelligence runs in the earbud. What does run there, and why it physically has to, is every constraint this blog writes about, compressed into five grams sitting in your ear canal.
The earbud is the smallest federation you own.
What runs in the bud
Four jobs that have to be available continuously, and only one of them is a neural network. Not everything that looks like intelligence at the edge is a learned model, and the earbud is where that becomes obvious.
- Noise cancelling. An adaptive filter in dedicated silicon, not a network.* It cannot be offloaded: sound covers the last three centimetres of ear canal in about ninety microseconds. Neither vendor publishes the loop rate outright, but both quote figures in that neighbourhood. Apple says the H2 chip, in its hearing protection feature, reduces loud, intermittent noise 48,000 times per second, which is once every twenty-one microseconds. Google says the Tensor A1 lets Pixel Buds Pro 2 adapt to the environment up to 3 million times per second.
- Wake word. A small classifier scoring every short window of audio against one fixed phrase, Hey Siri or OK Google, and nothing else. It is the only model running while the bud is idle, and it is the switch for the expensive voice path: the assistant pipeline on the phone, and anything beyond it, stays asleep until it fires. Its whole job is to say no, cheaply, millions of times for every once it says yes, which is why it has to be tiny.
- Voice activity. An accelerometer picks up your own voice through your jaw, which is how a bud tells your speech from the speech of whoever is standing beside you.
- In-ear detection. An infrared emitter with a photodiode beside it, or a capacitive plate, sampling one number many times a second: light returned, or a shift in capacitance. Skin reads differently from a pocket lining, and a threshold on that number is what pauses your music when you take a bud out. Some products run a small classifier over the same readings instead, to tell in-ear from in-pocket from half-seated.
Offloading that first loop is not slow. It is meaningless.
* The usual algorithm, filtered-x least mean squares, is one linear layer trained online by gradient descent against the residual measured at a microphone inside your ear. No pre-training, no nonlinearity, no depth. Signal processing would not call it AI, and mathematically it is a close cousin. The learned models sit above it at a human rate, choosing its settings rather than producing the anti-noise. Which is the lesson worth carrying: what a datasheet calls AI can be a very small adaptive filter.
Energy, and open models
A teardown of the original AirPods Pro puts one earbud’s cell at about 0.16 Wh, and Apple claims up to 8 hours with noise cancelling on for the current generation. A battery of that class over a session of that length puts the whole device, radio and codec and microphones included, in the tens of milliwatts.
The models in your ear are proprietary. The same jobs have open implementations you can put on a scale.
| the job | the model | parameters | on disk |
|---|---|---|---|
| voice activity detection | Silero VAD | about 309,000 | roughly 1.2 MB |
| keyword spotting | MLPerf Tiny’s DS-CNN | 24,908 | 52.5 KB, int8 |
| noise suppression on outgoing voice | RNNoise | about 60,000 | 85 KB, 8-bit weights |
All three are kilobytes to low megabytes, and all three work in millisecond frames. The two jobs with no row here are the two that are barely models: noise cancelling, where a microsecond deadline favours a filter over a network, and in-ear detection, where a comparator on one sensor reading usually does the job. RNNoise is the near miss, and a different job with a different budget: it cleans outgoing voice in millisecond frames, where the in-ear filter answers in microseconds.
What Kernwerk can help you with
The earbud is the smallest federation you own: four always-on jobs, each running where the physics and the battery allow. Almost everything you would call the AI is somewhere else, and the wake word detector is what decides when the assistant gets to hear you. Not every microphone path runs through it, calls and recordings do not, but the assistant one does: a privacy boundary implemented as an energy optimisation, shipped at scale because the battery demanded it rather than because anyone legislated it.
- Make it fit, and make it hold its deadline. Compress the model onto the part you already chose, and measure the tail rather than the mean, because a loop that is quick on average and late once in a thousand cycles is not a reflex.
- Make it cheap enough to leave running, and keep the data where it was measured. Size the model to the battery so the expensive tier stays asleep, stop the raw signal at the sensor, and sign what ships, because a model on a part an attacker can buy comes off the board unless something stops it.
The hardware figures here are vendor claims and teardowns, linked and dated to September 2026. The rest are derived from them: ninety microseconds is three centimetres at the speed of sound, twenty-one is the reciprocal of 48,000, and the milliwatts are a cell divided by a runtime.
If you have a model that has to run in a milliwatt-class device, that is the kind of deployment we work on. We are looking for design partners. Talk to us.