6 papers · 1 filter
FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention
Hangtian Zhao, Xiang Chen, Yizhe Li +3
In this paper we propose FastViDAR, a novel framework that takes four fisheye camera inputs and produces a full depth map along with per-camera depth, fusion depth, and…
Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
Mang Cao, Sanping Zhou, Yizhe Li +3
Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing exi…
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
Yizhe Li, Sanping Zhou, Zheng Qin +1
Dense video captioning is a challenging task that aims to localize and caption multiple events in an untrimmed video. Recent studies mainly follow the transformer-based architectur…
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
Yizhen Li, Dell Zhang, Xuelong Li +1
Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise visual perception and vision-te…
Advancing Pre-trained Teacher: Towards Robust Feature Discrepancy for Anomaly Detection
Canhui Tang, Sanping Zhou, Yizhe Li +2
With the wide application of knowledge distillation between an ImageNet pre-trained teacher model and a learnable student model, unsupervised anomaly detection has witnessed a sign…
Single-Shot and Multi-Shot Feature Learning for Multi-Object Tracking
Yizhe Li, Sanping Zhou, Zheng Qin +3
Multi-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a vi…