4 papers
Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics
Jiahua Wang, Leqi Zheng, Jialong Wu +2
World models simulate environmental dynamics to enable embodied agents to plan and reason about future states. While real-world perception is inherently multimodal, existing approa…
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
Xiaochen Yang, Hao Fang, Jiawei Kong +3
Although large vision-language models (LVLMs) have demonstrated remarkable capabilities, they are prone to hallucinations in multi-image tasks. We attribute this issue to limitatio…
What Papers Don't Tell You: Recovering Tacit Knowledge for Automated Paper Reproduction
Lehui Li, Ruining Wang, Haochen Song +8
Automated paper reproduction -- generating executable code from academic papers -- is bottlenecked not by information retrieval but by the tacit knowledge that papers inevitably le…
What Should I Cite? A RAG Benchmark for Academic Citation Prediction
Leqi Zheng, Jiajun Zhang, Canzhi Chen +13
With the rapid growth of Web-based academic publications, more and more papers are being published annually, making it increasingly difficult to find relevant prior work. Citation…