3 papers
cs.CV2025
DMC: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
Jiayi Zou, Chaofan Chen, Bing-Kun Bao +1
Egocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos…
cs.CV2025
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
Yuhang Hu, Zhenyu Yang, Shihan Wang +5
The rapid growth of streaming video applications demands multimodal models with enhanced capabilities for temporal dynamics understanding and complex reasoning. However, current Vi…
cs.CV2024
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
Ming Tao, Bing-Kun Bao, Hao Tang +2
Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trai…