4 papers · 1 filter
EEA: Exploration-Exploitation Agent for Long Video Understanding
Te Yang, Xiangyu Zhu, Bo Wang +3
Long-form video understanding requires efficient navigation of extensive visual data to pinpoint sparse yet critical information. Current approaches to longform video understanding…
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
Te Yang, Jian Jia, Xiangyu Zhu +9
Large Language Models (LLMs) have strong instruction-following capability to interpret and execute tasks as directed by human commands. Multimodal Large Language Models (MLLMs) hav…
Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
Shengyuan Liu, Bo Wang, Ye Ma +6
Existing subject-driven text-to-image generation models suffer from tedious fine-tuning steps and struggle to maintain both text-image alignment and subject fidelity. For generatin…
Knowledge Condensation and Reasoning for Knowledge-based VQA
Dongze Hao, Jian Jia, Longteng Guo +8
Knowledge-based visual question answering (KB-VQA) is a challenging task, which requires the model to leverage external knowledge for comprehending and answering questions grounded…