Model Detail
whisper-large-v3-turbo
—whisper-large-v3-turbo is an audio model with 404M parameters released by OpenAI. The model is registered under the automatic-speech-recognition pipeline tag on Hugging Face, distributed under the permissive mit license.
whisper-large-v3-turbo ships with 404M parameters. The mit license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
whisper-large-v3-turbo is best fit for speech recognition, transcription, or speech synthesis depending on the task head. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition
arXiv:2609.01737v1 Announce Type: new Abstract: Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users. This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a controlled study of domai
Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models
arXiv:2609.01723v1 Announce Type: cross Abstract: Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content. Existing black-box m
Context-Aware Interleaved Batching for WhisperX
arXiv:2608.31170v1 Announce Type: new Abstract: While WhisperX accelerates speech transcription via intra-audio batching, it isolates audio segments, losing the historical context needed for coherent punctuation and terminology transcription. Conversely, standard Whisper retains context sequentially
Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper
arXiv:2608.30927v1 Announce Type: cross Abstract: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a
Concepts Whisper: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations
arXiv:2605.01609v2 Announce Type: replace-cross Abstract: We find that transformer concept representations systematically anti-concentrate in the spectral tail of the unembedding covariance, encoding word-level concepts in low-variance directions across a 17-model core suite and an expanded set of 2
Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
arXiv:2608.26060v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through the use of large multilingual foundation models. However, most advances remain concentrated on high-resource languages, while indigenous langua