activity
20242026
collaborators

7 papers

cs.CV2026

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

Yifu Qiu, Yftah Ziser, Anna Korhonen +2

Can unified vision-language models (VLMs) perform forward dynamics prediction (FDP), i.e., predicting the future state (in image form) given the previous observation and an action…

cs.CV2026

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation

Ken Deng, Yifu Qiu, Yoni Kasten +2

We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view spatial reasoning. We cast RCPE as a d…

cs.LG2026

Self-Improving World Modelling with Latent Actions

Yifu Qiu, Zheng Zhao, Waylon Li +4

Internal modelling of the world -- predicting transitions between previous states and next states under actions -- is essential to reasoning and planning for LLMs and V…

cs.CL2025

Iterative Multilingual Spectral Attribute Erasure

Shun Shao, Yftah Ziser, Zheng Zhao +3

Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between langu…

cs.CL2025

Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models

Yifu Qiu, Varun Embar, Yizhe Zhang +3

Recent advancements in long-context language models (LCLMs) promise to transform Retrieval-Augmented Generation (RAG) by simplifying pipelines. With their expanded context windows,…

cs.CL2025

What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations

Dongqi Liu, Chenxi Whitehouse, Xi Yu +6

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed…