Model Detail
Voxtral-Mini-4B-Realtime-2602
—Voxtral-Mini-4B-Realtime-2602 is an audio model with 4B parameters released by Mistral. The model is registered under the automatic-speech-recognition pipeline tag on Hugging Face, distributed under the permissive apache-2.0 license.
Voxtral-Mini-4B-Realtime-2602 ships with 4B parameters. Total weight footprint is approximately 4.4 GB, which is the relevant figure when planning local-inference VRAM. The apache-2.0 license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
Downloads of Voxtral-Mini-4B-Realtime-2602 have moved -3.3% over the trailing seven days. That is a slight downtrend, consistent with normal cooling as newer models compete for the same workloads. These numbers are signal, not guarantee — week-over-week download counts on Hugging Face also reflect mirror traffic, CI scrapes, and one-off benchmarking runs.
Voxtral-Mini-4B-Realtime-2602 is best fit for speech recognition, transcription, or speech synthesis depending on the task head. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Voxtral Realtime
arXiv:2602.11298v3 Announce Type: replace Abstract: We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adapt offline models through chunking or sliding windows, Voxtral Realti
Voxtral TTS
arXiv:2603.25551v2 Announce Type: replace Abstract: We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid architecture that combines auto-regressive generation of semantic sp