activity
20242026
most citedResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding

4 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2026

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

Minghang Zheng, Zihao Yin, Yi Yang +2

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causi…

cs.CV2026

Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

Shengli Zhou, Minghang Zheng, Feng Zheng +1

Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limit…

cs.CV2025

InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects

Xinhao Cai, Minghang Zheng, Xin Jin +1

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient…

cs.CV2025

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

Minghang Zheng, Yuxin Peng, Benyuan Sun +2

In this paper, we tackle the task of online video temporal grounding (OnVTG), which requires the model to locate events related to a given text query within a video stream. Unlike…

cs.CV20244 cited

ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding

Minghang Zheng, Jiahua Zhang, Qingchao Chen +2

Visual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objec…

cs.CV2024

Training-free Video Temporal Grounding using Large-scale Pre-trained Models

Minghang Zheng, Xinhao Cai, Qingchao Chen +2

Video temporal grounding aims to identify video segments within untrimmed videos that are most relevant to a given natural language query. Existing video temporal localization mode…