Blog · September 29, 2026 · by Maxime Carriere

What is actually running inside your earbuds?

You talk to an AI through your earbuds every day. Almost none of the large-model intelligence runs in the earbud. What does run there, and why it physically has to, is every constraint this blog writes about, compressed into five grams sitting in your ear canal.

The earbud is the smallest federation you own.

An outline drawing of a pair of earbuds beside four named models, each with the constraint that keeps it on the bud. Noise cancelling, which computes the inverse of the noise about every 21 microseconds. Wake word, which listens for one phrase so that everything else can stay asleep. Voice activity, where an accelerometer feels your jaw so your voice is not a stranger's. In-ear detection, optical or capacitive, asking whether the bud is in an ear at all. A closing line notes that the whole device, radio and codec and microphones included, lives on tens of milliwatts.
Four models, none of them ever switched off, sharing a budget of tens of milliwatts. September 2026.

What runs in the bud

Four jobs that have to be available continuously, and only one of them is a neural network. Not everything that looks like intelligence at the edge is a learned model, and the earbud is where that becomes obvious.

  • Noise cancelling. An adaptive filter in dedicated silicon, not a network.* It cannot be offloaded: sound covers the last three centimetres of ear canal in about ninety microseconds. Neither vendor publishes the loop rate outright, but both quote figures in that neighbourhood. Apple says the H2 chip, in its hearing protection feature, reduces loud, intermittent noise 48,000 times per second, which is once every twenty-one microseconds. Google says the Tensor A1 lets Pixel Buds Pro 2 adapt to the environment up to 3 million times per second.
  • Wake word. A small classifier scoring every short window of audio against one fixed phrase, Hey Siri or OK Google, and nothing else. It is the only model running while the bud is idle, and it is the switch for the expensive voice path: the assistant pipeline on the phone, and anything beyond it, stays asleep until it fires. Its whole job is to say no, cheaply, millions of times for every once it says yes, which is why it has to be tiny.
  • Voice activity. An accelerometer picks up your own voice through your jaw, which is how a bud tells your speech from the speech of whoever is standing beside you.
  • In-ear detection. An infrared emitter with a photodiode beside it, or a capacitive plate, sampling one number many times a second: light returned, or a shift in capacitance. Skin reads differently from a pocket lining, and a threshold on that number is what pauses your music when you take a bud out. Some products run a small classifier over the same readings instead, to tell in-ear from in-pocket from half-seated.
Noise drawn as a wave arriving at a microphone, then a small model box, then an inverted anti-noise wave travelling on to an eardrum. A bracket under the whole path is annotated: sound covers the last 3 cm of canal in about 90 microseconds, and Apple says its chip answers 48,000 times a second. A dashed line drops from the model to a cloud icon struck through, labelled 30 to 100 ms to a server, the sound won long ago.
The 90 µs is 3 cm of ear canal at the speed of sound, 343 m/s. The 48,000 per second is Apple's published figure. September 2026.

Offloading that first loop is not slow. It is meaningless.

* The usual algorithm, filtered-x least mean squares, is one linear layer trained online by gradient descent against the residual measured at a microphone inside your ear. No pre-training, no nonlinearity, no depth. Signal processing would not call it AI, and mathematically it is a close cousin. The learned models sit above it at a human rate, choosing its settings rather than producing the anti-noise. Which is the lesson worth carrying: what a datasheet calls AI can be a very small adaptive filter.

Energy, and open models

A teardown of the original AirPods Pro puts one earbud’s cell at about 0.16 Wh, and Apple claims up to 8 hours with noise cancelling on for the current generation. A battery of that class over a session of that length puts the whole device, radio and codec and microphones included, in the tens of milliwatts.

Three bars to the same scale, showing battery energy in watt-hours. One earbud at 0.16 Wh from an iFixit teardown of the 2019 AirPods Pro, marked times one. A smartwatch at 1.4 Wh, an Apple Watch Series 11 in 46 mm, marked times nine. A phone at 14.4 Wh, an iPhone 17, marked times ninety.
One teardown of a 2019 earbud against two current published capacities. September 2026.

The models in your ear are proprietary. The same jobs have open implementations you can put on a scale.

the job the model parameters on disk
voice activity detection Silero VAD about 309,000 roughly 1.2 MB
keyword spotting MLPerf Tiny’s DS-CNN 24,908 52.5 KB, int8
noise suppression on outgoing voice RNNoise about 60,000 85 KB, 8-bit weights

All three are kilobytes to low megabytes, and all three work in millisecond frames. The two jobs with no row here are the two that are barely models: noise cancelling, where a microsecond deadline favours a filter over a network, and in-ear detection, where a comparator on one sensor reading usually does the job. RNNoise is the near miss, and a different job with a different budget: it cleans outgoing voice in millisecond frames, where the in-ear filter answers in microseconds.

What Kernwerk can help you with

The earbud is the smallest federation you own: four always-on jobs, each running where the physics and the battery allow. Almost everything you would call the AI is somewhere else, and the wake word detector is what decides when the assistant gets to hear you. Not every microphone path runs through it, calls and recordings do not, but the assistant one does: a privacy boundary implemented as an energy optimisation, shipped at scale because the battery demanded it rather than because anyone legislated it.

  • Make it fit, and make it hold its deadline. Compress the model onto the part you already chose, and measure the tail rather than the mean, because a loop that is quick on average and late once in a thousand cycles is not a reflex.
  • Make it cheap enough to leave running, and keep the data where it was measured. Size the model to the battery so the expensive tier stays asleep, stop the raw signal at the sensor, and sign what ships, because a model on a part an attacker can buy comes off the board unless something stops it.

The hardware figures here are vendor claims and teardowns, linked and dated to September 2026. The rest are derived from them: ninety microseconds is three centimetres at the speed of sound, twenty-one is the reciprocal of 48,000, and the milliwatts are a cell divided by a runtime.

If you have a model that has to run in a milliwatt-class device, that is the kind of deployment we work on. We are looking for design partners. Talk to us.