2 papers
cs.CV2026
Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment
Zihao Wang, Xi Xiang, Yuwen Sun +5
Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, tex…
cs.CV2026
PercepCap: Video Captioner with Structured Spatio-Temporal Perception
Yifan Xu, Zihao Wang, Zhixiao Wang +6
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occ…