3 papers
cs.CV2026
REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding
Boyang Li, Chenhui Gou, Jianfei Cai
Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A common zero-shot strategy asks a…
cs.CL2026
Sample-Efficient Learning from Agent Experience
Chenhui Gou, Haoqin Tu, Yunhao Fang +2
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offer…
cs.CV2026
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
Ziyu Ma, Shidong Yang, Yuxiang Ji +5
High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when the image is processed as a w…