8 papers
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan +4
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challengi…
OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam
Spatial reasoning is fundamental to auditory perception, yet current audio large language models (ALLMs) largely rely on unstructured binaural cues and single step inference. This…
LLaSA: A Sensor-Aware LLM for Natural Language Reasoning of Human Activity from IMU Data
Sheikh Asif Imran, Mohammad Nur Hossain Khan, Subrata Biswas +1
Wearable systems can recognize activities from IMU data but often fail to explain their underlying causes or contextual significance. To address this limitation, we introduce two l…
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam
Multimodal question answering (QA) often requires identifying which video, audio, or sensor tokens are relevant to the question. Yet modality disagreements are common: off-camera s…
Mindfulness Meditation and Respiration: Accelerometer-Based Respiration Rate and Mindfulness Progress Estimation to Enhance App Engagement and Mindfulness Skills
Mohammad Nur Hossain Khan, David creswell, Jordan Albert +5
Mindfulness training is widely recognized for its benefits in reducing depression, anxiety, and loneliness. With the rise of smartphone-based mindfulness apps, digital meditation h…
LOCUS: LOcalization with Channel Uncertainty and Sporadic Energy
Subrata Biswas, Mohammad Nur Hossain Khan, Violet Colwell +2
Accurate sound source localization (SSL), such as direction-of-arrival (DoA) estimation, relies on consistent multichannel data. However, batteryless systems often suffer from miss…