activity
20232025
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention

Hangtian Zhao, Xiang Chen, Yizhe Li +3

In this paper we propose FastViDAR, a novel framework that takes four fisheye camera inputs and produces a full depth map along with per-camera depth, fusion depth, and…

cs.CV2025

Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction

Mang Cao, Sanping Zhou, Yizhe Li +3

Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing exi…

cs.CV2025

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning

Yizhe Li, Sanping Zhou, Zheng Qin +1

Dense video captioning is a challenging task that aims to localize and caption multiple events in an untrimmed video. Recent studies mainly follow the transformer-based architectur…

cs.CV2025

Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations

Yizhen Li, Dell Zhang, Xuelong Li +1

Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise visual perception and vision-te…

cs.CV2024

Advancing Pre-trained Teacher: Towards Robust Feature Discrepancy for Anomaly Detection

Canhui Tang, Sanping Zhou, Yizhe Li +2

With the wide application of knowledge distillation between an ImageNet pre-trained teacher model and a learnable student model, unsupervised anomaly detection has witnessed a sign…

cs.CV2023

Single-Shot and Multi-Shot Feature Learning for Multi-Object Tracking

Yizhe Li, Sanping Zhou, Zheng Qin +3

Multi-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a vi…