4 papers
DAVSS: Distilled Audio-Visual State Space Models
Saurabhchand Bhati, Mrudula Athi, Amit S. Chhetri +1
State-space models (SSMs) distilled from transformer teachers combine the performance of transformers with the efficiency of SSMs. We extend the Transformer-SSM knowledge distillat…
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati +4
Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (S…
USAD: Universal Speech and Audio Representation via Distillation
Heng-Jui Chang, Saurabhchand Bhati, James Glass +1
Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we pr…
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Heng-Jui Chang, Hongyu Gong, Changhan Wang +2
Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous sp…