2 papers
cs.LG2026
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen +8
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multi…
cs.CV2026
Reliability-Prioritized Fine-Grained Generation in Multimodal Large
Xiaomeng Fan, Wei Wu, Yuwei Wu +9
Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generati…