collaborators

26 papers

cs.SD2026

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Yanqiu Li, Yang Xiao, Jisheng Bai +3

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which…

cs.RO2026

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding

Zhewei Zhang, Puyue Wang, Guanren Qiao +10

Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…

cs.AI2026

Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment

Xiangyuan Xue, Yang Yu, Yan Gao +5

Pulmonary embolism (PE) is a high risk cardiopulmonary condition whose management requires both timely diagnosis and reliable assessment of future clinical risk. Because PE care ro…

cs.CV2026

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Kai Bian, Xucheng Guo, Bin Chen +4

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widesprea…

cs.LG2026

Localizing and Editing Knowledge in Large Audio-Language Models

Sung Kyun Chung, Jiaheng Dong, Qiuchi Hu +3

Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual information. Yet they are trained…

cs.AI2026

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

Di Zhu, Yu Yvonne Wu, Hong Jia +3

Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines o…