6 papers
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen +8
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multi…
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
Niccolo Avogaro, Nayanika Debnath, Li Mi +6
Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for vision-language models (VLMs). Unstruc…
Knowledge-aware Visual Question Generation for Remote Sensing Images
Siran Li, Li Mi, Javiera Castillo-Navarro +1
With the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific information or performing image retriev…
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
Siran Li, Li Mi, Javiera Castillo-Navarro +1
With the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific information or performing semantic imag…
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
Li Mi, Manon Bechaz, Zeming Chen +2
Active Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search ar…
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
Silin Gao, Sheryl Mathew, Li Mi +6
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to…