activity
20242026
collaborators

13 papers

cs.CV2026

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Andong Lu, Ziyi Zha, Jiandong Jin +4

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attemp…

cs.CV2026

Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark

Qishun Wang, Yapeng Li, Bin Luo +2

RGB-Thermal (RGBT) Video Object Detection (VOD) has gained significant traction due to its ability to overcome the limitations of conventional RGB-based VOD under challenging condi…

eess.IV2026

Compositional-Degradation UAV Image Restoration: Conditional Decoupled MoE Network and A Benchmark

Jinquan Yan, Zhicheng Zhao, Zhengzheng Tu +3

UAV images are critical for applications such as large-area mapping, infrastructure inspection, and emergency response. However, in real-world flight environments, a single image i…

cs.CV2025

Vehicle-centric Perception via Multimodal Structured Pre-training

Wentao Wu, Xiao Wang, Chenglong Li +2

Vehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existin…

cs.CV2025

ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification

Shihao Li, Chenglong Li, Aihua Zheng +2

Multi-spectral object re-identification (ReID) brings a new perception perspective for smart city and intelligent transportation applications, effectively addressing challenges fro…

cs.CV2025

CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework

Wentao Wu, Xiao Wang, Chenglong Li +4

Event cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. S…