7 papers
Backbone-Agnostic Stochastic Perturbation Learning for End-to-End Real-World Image Dehazing
Bingcai Wei, Yuning Cui, Mingyu Liu +5
Real-world paired image dehazing remains challenging because haze degradation is spatially non-uniform, illumination-dependent, and physically ambiguous even when haze-free referen…
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Xingcheng Zhou, Hao Guo, Rui Song +5
Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses unde…
LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results
Xiang Chen, Hao Li, Jiangxin Dong +54
This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world all-in-one image restoration…
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
Xingcheng Zhou, Mingyu Liu, Walter Zimmer +2
We present Scene-Graph Based Multi-Modal Traffic Agent (SGTA), a modular framework for traffic video understanding that combines structured scene graphs with multi-modal reasoning.…
CAS-IQA: Teaching Vision-Language Models for Synthetic Angiography Quality Assessment
Bo Wang, De-Xing Huang, Xiao-Hu Zhou +5
Synthetic X-ray angiographies generated by modern generative models hold great potential to reduce the use of contrast agents in vascular interventional procedures. However, low-qu…
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
Xingcheng Zhou, Konstantinos Larintzakis, Hao Guo +7
We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 vid…