2 papers
cs.CV2026
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
Chenghua Zhu, Zhaolu Kang, Qifan Shi +8
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sam…
cs.MM2026
When Generated Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception
Guangyuan Dong, Chuang Liu, Yangchen Zeng +6
Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retri…