arxiv
PublishedJune 10, 2026 at 4:00 AM
Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
Publisher summary· verbatim
arXiv:2412.11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture. There has been a huge surge
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivLearning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance6harxivEvaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability6harxivSelf-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems6harxivAre we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs6hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗