51 citations · 51 across the 3 of their papers we have counts for
4 papers
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
Yiming Zhao, Yu Zeng, Wenxuan Huang +11
Large Vision-Language Models (LVLMs) have shown significant progress in video understanding, yet they face substantial challenges in tasks requiring precise spatiotemporal localiza…
Focus on BEV: Self-calibrated Cycle View Transformation for Monocular Birds-Eye-View Segmentation
Jiawei Zhao, Qixing Jiang, Xuede Li +1
Birds-Eye-View (BEV) segmentation aims to establish a spatial mapping from the perspective view to the top view and estimate the semantic maps from monocular images. Recent studies…
Transformer-based Dual Relation Graph for Multi-label Image Recognition
Jiawei Zhao, Ke Yan, Yifan Zhao +3
The simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsist…
RGB-D Salient Object Detection with Ubiquitous Target Awareness
Yifan Zhao, Jiawei Zhao, Jia Li +1
Conventional RGB-D salient object detection methods aim to leverage depth as complementary information to find the salient regions in both modalities. However, the salient object d…