3 papers
cs.CV2026
Video Generation Models are General-Purpose Vision Learners
Letian Wang, Chuhan Zhang, Rishabh Kabra +9
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a genera…
cs.CV2025
VoCap: Video Object Captioning and Segmentation from Any Prompt
Jasper Uijlings, Xingyi Zhou, Xiuye Gu +5
Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose…
cs.CV2024
HAMMR: HierArchical MultiModal React agents for generic VQA
Lluis Castrejon, Thomas Mensink, Howard Zhou +3
Combining Large Language Models (LLMs) with external specialized tools (LLMs+tools) is a recent paradigm to solve multimodal tasks such as Visual Question Answering (VQA). While th…