Our Signal Board has flagged this model four times without writing it up. It’s now had five consecutive days of likes growth — 168 on 24 September, 1,359 today, up 40% in the last day alone. That pattern earned Ming-Image a piece yesterday, and it earns this one today.
Edge0’s Audio8-ASR-Infinite is a streaming speech recognition model, and two of its properties are unusually relevant to interactive work.
The specifications
- 4 billion parameters, bfloat16, 8.17 GB
- Chinese and English
- Native streaming, making 12.5 decoding decisions per second at selectable intervals
- Apache 2.0 — a genuinely permissive licence, which after the last fortnight of research-only and ambiguous terms is worth noting on its own
- A preview release, described as focused on transcription foundations
Architecture is a pairing rather than a single model: a Voxtral Realtime audio tower (32 layers, 1280 hidden) feeding a Qwen2.5-3B-Instruct decoder (36 layers, 2048 hidden), plus semantic VAD.
”Infinite” means constant memory, not infinite context
The name is doing real work. Audio8 handles unlimited-length audio via a rolling KV cache, which keeps both memory and latency constant, even in 24/7 operation.
That’s the property most transcription systems don’t have and most installations need.
A conventional streaming ASR either accumulates context — so memory and latency grow with session length until something dies — or it chunks, restarting periodically and losing continuity at every seam. Neither survives being switched on in a gallery in October and left running until January.
A rolling cache bounds the problem: the model keeps a fixed window of recent context and discards the rest. You lose the ability to reference something said an hour ago, which for transcription is a reasonable trade, and you gain a process whose resource use on day ninety looks like day one.
The semantic VAD is the more interesting half
Standard voice activity detection answers one question: is there speech energy right now? It’s a threshold, and it’s why voice interfaces interrupt people.
Audio8’s semantic VAD distinguishes thinking pauses, stuttering, and real end of turn.
Those are three completely different silences that sound similar to an energy detector. A speaker gathering a thought, a speaker stumbling mid-word, and a speaker who has finished all produce a gap — and a system that treats them identically will cut off anyone who hesitates.
This lands one day after we covered a paper on backchannels in full-duplex models, which found that a model’s hidden states already anticipate human turn-taking timing. The two results point the same direction: the interesting problem in speech interfaces is no longer transcription, it’s timing. Getting the words right is largely solved. Knowing when someone has finished is not, and it’s what separates a system that feels attentive from one that feels rude.
For installation and performance work
Three concrete reasons to care:
Long-running voice pieces become viable. Constant memory and latency is the difference between a demo and an exhibition.
Hesitation stops breaking the interaction. Gallery visitors talking to an artwork are uncertain, halting, and often mid-thought. A VAD that tolerates that is worth more than a percentage point of word accuracy.
Apache 2.0 and 8.17 GB means you can actually deploy it — locally, commercially, without an API dependency or a per-minute bill for a piece that listens for three months.
The honest caveats: Chinese and English only, and it’s a preview release. Don’t build a commission around it without testing your own material first.