3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2025
Towards Text-Image Interleaved Retrieval
Xin Zhang, Ziqi Dai, Yongqi Li +7
Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved cont…
cs.AI2025★ 3 cited
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
Hao Fei, Shengqiong Wu, Wei Ji +4
Existing research of video understanding still struggles to achieve in-depth comprehension and reasoning in complex videos, primarily due to the under-exploration of two key bottle…