26 papers
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
Yanqiu Li, Yang Xiao, Jisheng Bai +3
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which…
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
Zhewei Zhang, Puyue Wang, Guanren Qiao +10
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…
Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment
Xiangyuan Xue, Yang Yu, Yan Gao +5
Pulmonary embolism (PE) is a high risk cardiopulmonary condition whose management requires both timely diagnosis and reliable assessment of future clinical risk. Because PE care ro…
Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models
Kai Bian, Xucheng Guo, Bin Chen +4
Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widesprea…
Localizing and Editing Knowledge in Large Audio-Language Models
Sung Kyun Chung, Jiaheng Dong, Qiuchi Hu +3
Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual information. Yet they are trained…
VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data
Di Zhu, Yu Yvonne Wu, Hong Jia +3
Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines o…