3 papers
cs.CV2026
IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
Jiapeng Li, Ping Wei, Wenjuan Han +2
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark mat…
cs.CV2023
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
Zhi Gao, Yuntao Du, Xintong Zhang +4
Utilizing large language models (LLMs) to compose off-the-shelf visual tools represents a promising avenue of research for developing robust visual assistants capable of addressing…
cs.CL2023
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Haozhe Zhao, Zefan Cai, Shuzheng Si +7
Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utiliz…