Skip to content

[ commentary ]

Whistle transcribes speech from a 16.9 MB file

Cactus Compute's Whistle is a 16.9 MB open speech recognition model that runs on CPU in the same C++ engine as its Needle model.

Published · on cactuscompute.com · 2 min read

Cactus Compute's Whistle is a 16.9 MB open speech recognition model that runs on CPU in the same C++ engine as its Needle model.

Cactus Compute has released Whistle, an open speech recognition model that ships as a single 16.9 MB file and runs on the CPU with no dependencies. The company announced it on 2 October 2026 on cactuscompute.com. The size is the point: speech recognition has usually meant a model measured in hundreds of megabytes, which decides where it can run at all.

Eight blocks, one file, seventeen targets

Whistle takes 16 kHz mono audio, up to 30 seconds in one pass, and transcribes English, German, French, Spanish, Italian, Dutch and Polish. The language is detected unless you name it. It does three jobs on the device: transcription, word timestamps with a start, end and probability per word, and a speech embedding taken from the encoder output, one row per 80 ms frame, without decoding a transcript.

The encoder is eight Simple Attention blocks, four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network — the same blocks Cactus Compute's text model Needle uses. The decoder is eight Laddered Simple Attention blocks at width 512, with gated cross attention to the encoder. The ladder matters for deployment: every depth from two layers up was trained as a model of its own, and a flag selects one at load time, while the encoder always runs all eight blocks.

The engine ships prebuilt for seventeen targets, from macOS and Linux through Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. Each folder holds a binary, a static library and a header, and loads any model container handed to it. The same binary can run speech, text, or both: given a clip and a tool list, it transcribes the audio, answers the transcript against the tools and returns one JSON object with the calls and the speech fields.

What the published numbers say

Cactus Compute reports 16.9 MB against 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. On ten seconds of audio on an Apple M4 Pro CPU, it reports 11.1 ms to the first token against 73.2 ms and 22.8 ms, and 1,319 decode tokens per second against 266 and 262. Word error rates are scored with the Whisper normalizers over 86,174 utterances for Whistle; the Whisper and Moonshine figures are the ones their authors published. Whistle is ahead on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average; Whisper base is ahead on TED-LIUM, AMI and the MLS average. The weights are on Hugging Face, and the engine source is on GitHub.

Source: cactuscompute.com — BARGO’s commentary on the linked source.

Source published: · Event date:

Which task costs your team the most hours every week? Tell us, and we will tell you whether AI can take it and what it would cost.

← all posts · RSS