7 papers
Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics
Yifu Qiu, Yftah Ziser, Anna Korhonen +2
Can unified vision-language models (VLMs) perform forward dynamics prediction (FDP), i.e., predicting the future state (in image form) given the previous observation and an action…
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
Ken Deng, Yifu Qiu, Yoni Kasten +2
We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view spatial reasoning. We cast RCPE as a d…
Self-Improving World Modelling with Latent Actions
Yifu Qiu, Zheng Zhao, Waylon Li +4
Internal modelling of the world -- predicting transitions between previous states and next states under actions -- is essential to reasoning and planning for LLMs and V…
Iterative Multilingual Spectral Attribute Erasure
Shun Shao, Yftah Ziser, Zheng Zhao +3
Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between langu…
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
Yifu Qiu, Varun Embar, Yizhe Zhang +3
Recent advancements in long-context language models (LCLMs) promise to transform Retrieval-Augmented Generation (RAG) by simplifying pipelines. With their expanded context windows,…
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
Dongqi Liu, Chenxi Whitehouse, Xi Yu +6
Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed…