2 papers
cs.CV2026
LensWalk: Agentic Video Understanding by Planning How You See in Videos
Keliang Li, Yansong Li, Hongze Shen +3
The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understan…
cs.CV2024
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
Qibo Chen, Weizhong Jin, Jianyue Ge +7
Recent research on universal object detection aims to introduce language in a SoTA closed-set detector and then generalize the open-set concepts by constructing large-scale (text-r…