activity
20202026
most citedRegularizing Contrastive Predictive Coding for Speech Applications

1 citations · 3 across the 18 of their papers we have counts for

collaborators
Showing eess.ASShow all

11 papers · 1 filter

eess.AS2026

DAVSS: Distilled Audio-Visual State Space Models

Saurabhchand Bhati, Mrudula Athi, Amit S. Chhetri +1

State-space models (SSMs) distilled from transformer teachers combine the performance of transformers with the efficiency of SSMs. We extend the Transformer-SSM knowledge distillat…

eess.AS2026

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati +4

Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (S…

eess.AS2025

Towards Audio Token Compression in Large Audio Language Models

Saurabhchand Bhati, Samuel Thomas, Hilde Kuehne +2

Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.g., 25 tokens/s), makin…

eess.AS2025

Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?

Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo +4

We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new St…

eess.AS2024★ 1 cited

State-Space Large Audio Language Models

Saurabhchand Bhati, Yuan Gong, Leonid Karlinsky +3

Large Audio Language Models (LALM) combine the audio perception models and the Large Language Models (LLM) and show a remarkable ability to reason about the input audio, infer the…

eess.AS2024

DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners

Saurabhchand Bhati, Yuan Gong, Leonid Karlinsky +3

State-space models (SSMs) have emerged as an alternative to Transformers for audio modeling due to their high computational efficiency with long inputs. While recent efforts on Aud…