collaborators

8 papers

cs.CV2026

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao +9

Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate r…

cs.LG2026

Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits

Andreas Grivas, Lorenzo Loconte, Emile van Krieken +6

Multi-token prediction (MTP) is a prominent strategy to significantly speed up generation in large language models (LLMs), especially in byte-level LLMs, which are tokeniser-free b…

cs.CL2026

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

Wai-Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang +1

Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dependent on curated prompts or…

cs.CV2026

Learning GUI Grounding with Spatial Reasoning from Visual Feedback

Yu Zhao, Wei-Ning Chen, Huseyin Atahan Inan +8

Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate on-screen coordinates for actions such…

cs.AI2026

Same Answer, Different Representations: Hidden instability in VLMs

Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena +6

The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions reflect stable multimodal processi…

cs.CV2025

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren +9

The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of ima…