8 papers
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving
Nan Yang, Zhanwen Liu, Linfeng Zhang +4
Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly i…
DPMambaIR: All-in-One Image Restoration via Degradation-Aware Prompt State Space Model
Zhanwen Liu, Sai Zhou, Yuchao Dai +3
All-in-One image restoration aims to address multiple image degradation problems using a single model, offering a more practical and versatile solution compared to designing dedica…
Bidirectional Image-Event Guided Fusion Framework for Low-Light Image Enhancement
Zhanwen Liu, Huanna Song, Yang Wang +3
Under extreme low-light conditions, frame-based cameras suffer from severe detail loss due to limited dynamic range. Recent studies have introduced event cameras for event-guided l…
PSTTS: A Plug-and-Play Token Selector for Efficient Event-based Spatio-temporal Representation Learning
Xiangmo Zhao, Nan Yang, Yang Wang +1
Mainstream event-based spatio-temporal representation learning methods typically process event streams by converting them into sequences of event frames, achieving remarkable perfo…
Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection
Nan Yang, Yang Wang, Zhanwen Liu +3
Existing RGB-Event detection methods process the low-information regions of both modalities (background in images and non-event regions in event data) uniformly during feature extr…
Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios
Zhanwen Liu, Yujing Sun, Yang Wang +3
The dynamic range limitation of conventional RGB cameras reduces global contrast and causes loss of high-frequency details such as textures and edges in complex traffic environment…