5 citations · 9 across the 9 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2025
Thinking While Listening: Simple Test Time Scaling For Audio Classification
Prateek Verma, Mert Pilanci
We propose a framework that enables neural models to "think while listening" to everyday sounds, thereby enhancing audio classification performance. Motivated by recent advances in…
cs.SD2024
Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
Prateek Verma
We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously…