2 citations · 2 across the 5 of their papers we have counts for
6 papers · 1 filter
APPO: Attention-guided Perception Policy Optimization for Video Reasoning
Henghui Du, Chang Zhou, Xi Chen +1
Complex video reasoning, actually, relies excessively on fine-grained perception rather than on expert (e.g., Ph.D, Science)-level reasoning. Through extensive empirical observatio…
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
Henghui Du, Chunjie Zhang, Xi Chen +2
Long Video Question-Answering (LVQA) presents a significant challenge for Multi-modal Large Language Models (MLLMs) due to immense context and overloaded information, which could a…
Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception
Ruotian Peng, Haiying He, Yake Wei +2
High-quality image captions play a crucial role in improving the performance of cross-modal applications such as text-to-image generation, text-to-video generation, and text-image…
Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Henghui Du, Guangyao Li, Chang Zhou +3
In recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal local…
On-the-fly Modulation for Balanced Multimodal Learning
Yake Wei, Di Hu, Henghui Du +1
Multimodal learning is expected to boost model performance by integrating information from different modalities. However, its potential is not fully exploited because the widely-us…
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
Guangyao Li, Henghui Du, Di Hu
The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal vide…