collaborators

16 papers

cs.SD2026

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia +1

Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithf…

cs.CL2026

RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

Ram Narayanan, Harshit Rajgarhia, Abhishek Mukherji

Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet en…

cs.AI2026

GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings

Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin +1

Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated conte…

cs.SD2026

DuplexWorld: Can voice agents help you get through the day?

Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli +3

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversatio…

cs.SD2026

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji +1

Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered…

eess.AS2026

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds

Harshit Rajgarhia, Shuubham Ojha, Akhil Pothanapalli +4

Medical acoustic signals such as respiratory sounds, cardiac auscultations, and cough audio carry rich diagnostic information, yet no existing benchmark evaluates multimodal reason…