4 papers
REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding
Boyang Li, Chenhui Gou, Jianfei Cai
Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A common zero-shot strategy asks a…
Sample-Efficient Learning from Agent Experience
Chenhui Gou, Haoqin Tu, Yunhao Fang +2
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offer…
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
Ziyu Ma, Shidong Yang, Yuxiang Ji +5
High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when the image is processed as a w…
DrVideo: Document Retrieval Based Long Video Understanding
Ziyu Ma, Chenhui Gou, Hengcan Shi +4
Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The in…