5 papers
OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam
Spatial reasoning is fundamental to auditory perception, yet current audio large language models (ALLMs) largely rely on unstructured binaural cues and single step inference. This…
Mindfulness Meditation and Respiration: Accelerometer-Based Respiration Rate and Mindfulness Progress Estimation to Enhance App Engagement and Mindfulness Skills
Mohammad Nur Hossain Khan, David creswell, Jordan Albert +5
Mindfulness training is widely recognized for its benefits in reducing depression, anxiety, and loneliness. With the rise of smartphone-based mindfulness apps, digital meditation h…
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
Ashe Neth, Sawinder kaur, Mohammad Nur Hossain Khan +3
Existing pruning methods are typically applied during training or compile time and often rely on structured sparsity. While compatible with low-power microcontrollers (MCUs), struc…
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam
Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quan…
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam
Multimodal question answering (QA) often requires identifying which video, audio, or sensor tokens are relevant to the question. Yet modality disagreements are common: off-camera s…