2 papers
cs.CV2024
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
Yian Li, Wentao Tian, Yang Jiao +5
Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and ext…
cs.CV2023
Open-Vocabulary Video Relation Extraction
Wentao Tian, Zheng Wang, Yuqian Fu +2
A comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions. However, many current video understanding tasks pr…