162 citations · 200 across the 13 of their papers we have counts for
18 papers
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
Jiahao Nie, Wenbin An, Gongjie Zhang +4
Despite recent advances in Video Large Language Models (Vid-LLMs), Temporal Video Grounding (TVG), which aims to precisely localize time segments corresponding to query events, rem…
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
Jiahao Nie, Gongjie Zhang, Wenbin An +4
Though Multi-modal Large Language Models (MLLMs) have recently achieved significant progress, they often struggle to understand diverse and complicated inter-object relations. Spec…
Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining
Jiahao Nie, Yun Xing, Gongjie Zhang +5
Cross-Domain Few-Shot Segmentation (CD-FSS) poses the challenge of segmenting novel categories from a distinct domain using only limited exemplars. In this paper, we undertake a co…
Online Map Vectorization for Autonomous Driving: A Rasterization Perspective
Gongjie Zhang, Jiahao Lin, Shuang Wu +5
Vectorized high-definition (HD) map is essential for autonomous driving, providing detailed and precise environmental information for advanced perception and planning. However, cur…
Modeling Continuous Motion for 3D Point Cloud Object Tracking
Zhipeng Luo, Gongjie Zhang, Changqing Zhou +4
The task of 3D single object tracking (SOT) with LiDAR point clouds is crucial for various applications, such as autonomous driving and robotics. However, existing approaches have…
DETR4D: Direct Multi-View 3D Object Detection with Sparse Attention
Zhipeng Luo, Changqing Zhou, Gongjie Zhang +1
3D object detection with surround-view images is an essential task for autonomous driving. In this work, we propose DETR4D, a Transformer-based framework that explores sparse atten…