3 papers
cs.SD2026
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek +3
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained…
cs.SD2025
Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
Artem Dementyev, Chandan K. A. Reddy, Scott Wisdom +3
Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrain…
cs.HC2025
SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
Artem Dementyev, Dimitri Kanevsky, Samuel J. Yang +3
Speech-to-text capabilities on mobile devices have proven helpful for hearing and speech accessibility, language translation, note-taking, and meeting transcripts. However, our fou…