TL;DR
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Cactus Compute has released Whistle, a 16.9 MB speech-recognition model designed to run locally on CPUs without external dependencies. The company reports support for seven languages and 11.1 ms time to first token on an Apple M4 Pro, but the published performance figures are company benchmarks and the supplied report does not include full test results for every comparison.
Cactus Compute released Whistle, a speech-recognition model packaged as a 16.9 MB file for local use on devices including phones, wearables, robots and cars. The company says it transcribes speech in seven languages on a CPU, without external dependencies, and can return word-level timestamps and speech embeddings as well as text.
Whistle accepts 16 kHz mono audio clips of up to 30 seconds in English, German, French, Spanish, Italian, Dutch and Polish. Cactus Compute says it detects the language automatically unless the user specifies one. Its browser demonstration downloads the model on the first use and processes audio on the device; the company says the audio does not leave the device.
The model also produces start and end times for each transcribed word, along with a probability value. Its embedding mode returns a representation of the audio at 80-millisecond intervals without generating a transcript. A keyword-biasing option can give selected phrases more weight during decoding, according to the technical report.
Cactus reports an 11.1 ms time to first token for 10 seconds of audio on an Apple M4 Pro CPU. In the same test, it reports 1,319 decoded tokens per second, compared with 266 for Whisper base and 262 for Moonshine tiny v2. The company says it tested each system on its official runtime at default settings; Whistle used five-beam decoding. These are vendor-reported measurements, not an independent evaluation.
Small Models for Local Speech
A model that fits in 16.9 MB could make speech recognition easier to add to products with limited storage or unreliable internet access. Local processing can also reduce the need to send recordings to a remote service, although privacy depends on how the surrounding application handles audio and transcripts.
The proposed use cases span mobile devices and embedded systems, where power, memory and response time can constrain larger recognition systems. Cactus says Whistle uses the same C++ engine and model infrastructure as its Needle model, which could let developers combine speech recognition with other on-device processing. The source does not provide independent deployment results across the listed device categories.
on-device speech recognition software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Whistle Is Built
Cactus describes Whistle as an open speech-recognition model running in the same CPU engine as Needle. Its audio front end converts sound into log-mel features, then reduces the input to 375 frames for a 30-second clip. The model has an eight-block encoder and an eight-block decoder, with the decoder using cross-attention to read the encoded audio.
The company says developers can choose a decoder depth at load time, with trained options beginning at two layers; the encoder still runs all eight blocks. It also describes a silence check that measures a clip before decoding and returns an empty transcript below a loudness threshold. The model’s transcript output is capped at 320 tokens, according to the report.
Cactus compares Whistle’s word error rates with Whisper base and Moonshine tiny v2 across several speech datasets. It says Whistle scores better on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average, while Whisper base scores better on TED-LIUM, AMI and the MLS average. The report cautions that some datasets are unavailable for certain systems and that Whisper’s cited AMI result uses a different subset.
“one 16.9 MB file, runs on the CPU with no dependencies”
— Cactus Compute, in its October 2 release
As an affiliate, we earn on qualifying purchases.
Independent Results Still Needed
The published material is a vendor report, and the supplied benchmark description does not include the full word-error-rate values or detailed scoring tables. It is therefore not possible from this material alone to assess the size of the reported accuracy differences or reproduce the comparisons.
Performance figures are tied to an Apple M4 Pro CPU and the stated runtimes and settings. Results may differ on phones, microcontrollers and other supported target devices; the report does not provide power consumption, memory use during inference, or measurements across those platforms. The exact terms of the model’s openness and licensing are also not specified in the supplied release.
Cactus says audio stays on-device in its browser demo, but the source does not describe independent privacy testing or how applications built with Whistle may store transcripts. The company’s claim should not be taken as a guarantee about every integration.
privacy-focused speech transcription device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deployment and Testing Ahead
The release makes Whistle available for developers to try, including through a browser demonstration that downloads the model on first use. The immediate next step for adopters is to test recognition accuracy and latency on their own audio, languages and target hardware rather than assume the company’s benchmark results will carry over.
Independent testing could clarify how Whistle performs in noisy settings, with accents and specialist vocabulary, and how its reported speed compares when energy use and full application workloads are measured. Cactus has not specified a date for additional benchmarks or a broader release update in the material provided.
small speech recognition model for mobile
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model from Cactus Compute, released as a 16.9 MB file for local CPU use. It can produce transcripts, word timestamps and speech embeddings.
Which languages does it support?
Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. The company says Whistle detects the language automatically unless a user names it.
Does Whistle send recordings to a server?
Cactus says its browser demo processes audio on the device and that audio does not leave the device. The supplied report does not include an independent privacy audit or establish how every third-party integration handles audio.
How fast is Whistle?
For 10 seconds of audio on an Apple M4 Pro CPU, Cactus reports 11.1 ms to first token and 1,319 decoded tokens per second. Those are company benchmarks on specified hardware, not guarantees for other devices.
Has Whistle been shown to outperform Whisper on every test?
No. Cactus reports Whistle ahead on some listed datasets and Whisper base ahead on TED-LIUM, AMI and the MLS average. The company also notes that some comparisons are unavailable or use different dataset subsets.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
