6 papers
Unmixing The Crowd: Learning Persistent Speaker Representations from Mixture-Derived Multi-Speaker Embeddings
Sidharth Sidharth, Meysam Asgari, Hao-Wen Dong +1
We study whether persistent conversational speaker structure can be extracted directly from local overlapping speech mixtures. We propose a teacher-student framework that learns mi…
Sona: Real-Time Multi-Target Sound Attenuation for Noise Sensitivity
Jeremy Zhengqi Huang, Emani Hicks, Sidharth +2
For people with noise sensitivity, everyday soundscapes can be overwhelming. Existing tools such as active noise cancellation reduce discomfort by suppressing the entire acoustic e…
RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People
Xinyun Cao, Kexin Phyllis Ju, Chenglin Li +2
As virtual 3D environments become more prevalent, equitable access is essential for blind and low-vision (BLV) users, who face challenges with spatial awareness, navigation, and in…
EvolveCaptions: Empowering DHH Users Through Real-Time Collaborative Captioning
Liang-Yuan Wu, Dhruv Jain
Automatic Speech Recognition (ASR) systems often fail to accurately transcribe speech from Deaf and Hard of Hearing (DHH) individuals, especially during real-time conversations. Ex…
CapTune: Adapting Non-Speech Captions With Anchored Generative Models
Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu +1
Non-speech captions are essential to the video experience of deaf and hard of hearing (DHH) viewers, yet conventional approaches often overlook the diversity of their preferences.…
SonoCraftAR: Towards Supporting Personalized Authoring of Sound-Reactive AR Interfaces by Deaf and Hard of Hearing Users
Jaewook Lee, Davin Win Kyi, Leejun Kim +5
Augmented reality (AR) has shown promise for supporting Deaf and hard-of-hearing (DHH) individuals by captioning speech and visualizing environmental sounds, yet existing systems d…