1 citations · 3 across the 11 of their papers we have counts for
4 papers · 1 filter
Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs
Xi Xiao, Chen Liu, Chih-Ting Liao +9
Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text. Despite inheriting strong reason…
LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection
Mitchell Piehl, Muchao Ye
Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning ability and natural language-base…
Understanding Real-World Traffic Safety through RoadSafe365 Benchmark
Xinyu Liu, Darryl C. Jacob, Yuxin Liu +4
Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we in…
SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning
Zihao Zhao, Shengting Cao, Muchao Ye
Multi-modal large language models (MLLMs) have demonstrated significant progress in reasoning capabilities and shown promising effectiveness in video anomaly understanding (VAU) ta…