11 papers
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
Jiajun Wu, Haoyu Kang, Yining Sun +13
Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. Howeve…
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring
Yang Gao
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety…
Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment
Xiangyuan Xue, Yang Yu, Yan Gao +5
Pulmonary embolism (PE) is a high risk cardiopulmonary condition whose management requires both timely diagnosis and reliable assessment of future clinical risk. Because PE care ro…
Logics-Parsing-Omni Technical Report
Xin An, Jingyi Cai, Xiangyang Chen +22
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This fram…
HoRD: Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation
Puyue Wang, Jiawei Hu, Yan Gao +7
Humanoid robots can suffer significant performance drops under small changes in dynamics, task specifications, or environment setup. We propose HoRD, a two-stage learning framework…
Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
Xiangyuan Xue, Jiajun Lu, Yan Gao +3
Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains cha…