2 papers
cs.CV2026
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
Tingshu Mou, Jiabo He, Renying Wang +5
Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leavin…
cs.LG2025
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
Yige Li, Jiabo He, Hanxun Huang +3
Backdoor attacks have become a significant threat to the pre-training and deployment of deep neural networks (DNNs). Although numerous methods for detecting and mitigating backdoor…