collaborators

7 papers

cs.SD2026

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif +6

Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. Howeve…

cs.AI2026

MVG-KAN: Multi-View Geo-Wind Guided KAN for PM Forecasting

Cheng Huang, Muyao Guan, Jairus Yougui Railey +6

Accurate short-term PM forecasting is important for public health protection, air-quality early warning, and urban environmental management. However, PM variation i…

cs.SD2026

AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

Pengfei Zhang, Hoang H Nguyen, Yutong Song +6

Pathological speech from patients with neurodegenerative and neuromotor disorders is often acoustically distorted and linguistically fragmented, making pathological speech reconstr…

cs.AI2026

Can the Environment Speak for Itself? -GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

Yutong Song, Jiang Wu, Pengfei Zhang +4

Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics, such as patient distress a…

cs.CL2026

MedSpeak: A Knowledge Graph-Aided ASR Error Correction Framework for Spoken Medical QA

Yutong Song, Shiva Shrestha, Chenhan Lyu +5

Spoken question-answering (SQA) systems relying on automatic speech recognition (ASR) often struggle with accurately recognizing medical terminology. To this end, we propose MedSpe…

cs.MA2026

DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action Simulation

Yutong Song, Jiang Wu, Kazi Sharif +3

Simulating dementia patients with large language models (LLMs) is challenging due to the need to jointly model cognitive impairment, emotional dynamics, and nonverbal behaviors ove…