6 papers
Advancing Pre-trained Teacher: Towards Robust Feature Discrepancy for Anomaly Detection
Canhui Tang, Sanping Zhou, Yizhe Li +2
With the wide application of knowledge distillation between an ImageNet pre-trained teacher model and a learnable student model, unsupervised anomaly detection has witnessed a sign…
FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention
Hangtian Zhao, Xiang Chen, Yizhe Li +3
In this paper we propose FastViDAR, a novel framework that takes four fisheye camera inputs and produces a full depth map along with per-camera depth, fusion depth, and…
Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
Mang Cao, Sanping Zhou, Yizhe Li +3
Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing exi…
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
Yizhe Li, Sanping Zhou, Zheng Qin +1
Dense video captioning is a challenging task that aims to localize and caption multiple events in an untrimmed video. Recent studies mainly follow the transformer-based architectur…
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation
Ye Niu, Sanping Zhou, Yizhe Li +2
In many complex scenarios, robotic manipulation relies on generative models to estimate the distribution of multiple successful actions. As the diffusion model has better training…
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
Yizhen Li, Dell Zhang, Xuelong Li +1
Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise visual perception and vision-te…