7 papers
From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models
Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif +6
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. Howeve…
MVG-KAN: Multi-View Geo-Wind Guided KAN for PM Forecasting
Cheng Huang, Muyao Guan, Jairus Yougui Railey +6
Accurate short-term PM forecasting is important for public health protection, air-quality early warning, and urban environmental management. However, PM variation i…
AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction
Pengfei Zhang, Hoang H Nguyen, Yutong Song +6
Pathological speech from patients with neurodegenerative and neuromotor disorders is often acoustically distorted and linguistically fragmented, making pathological speech reconstr…
Can the Environment Speak for Itself? -GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents
Yutong Song, Jiang Wu, Pengfei Zhang +4
Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics, such as patient distress a…
MedSpeak: A Knowledge Graph-Aided ASR Error Correction Framework for Spoken Medical QA
Yutong Song, Shiva Shrestha, Chenhan Lyu +5
Spoken question-answering (SQA) systems relying on automatic speech recognition (ASR) often struggle with accurately recognizing medical terminology. To this end, we propose MedSpe…
DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action Simulation
Yutong Song, Jiang Wu, Kazi Sharif +3
Simulating dementia patients with large language models (LLMs) is challenging due to the need to jointly model cognitive impairment, emotional dynamics, and nonverbal behaviors ove…