activity
20242026
collaborators

9 papers

cs.CV2026

Learning GUI Grounding with Spatial Reasoning from Visual Feedback

Yu Zhao, Wei-Ning Chen, Huseyin Atahan Inan +8

Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate on-screen coordinates for actions such…

cs.CV2025

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren +9

The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of ima…

cs.CL2025

Noiser: Bounded Input Perturbations for Attributing Large Language Models

Mohammad Reza Ghasemi Madani, Aryo Pradipta Gema, Gabriele Sarti +3

Feature attribution (FA) methods are common post-hoc approaches that explain how Large Language Models (LLMs) make predictions. Accordingly, generating faithful attributions that r…

cs.CL2025

Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression

Nathan Godey, Alessio Devoto, Yu Zhao +4

Autoregressive language models rely on a Key-Value (KV) Cache, which avoids re-computing past hidden states during generation, making it faster. As model sizes and context lengths…

cs.CL2025

Structured Packing in LLM Training Improves Long Context Utilization

Konrad Staniszewski, Szymon Tworkowski, Sebastian Jaszczur +4

Recent advancements in long-context large language models have attracted significant attention, yet their practical applications often suffer from suboptimal context utilization. T…

cs.CL2025

Analysing the Residual Stream of Language Models Under Knowledge Conflicts

Yu Zhao, Xiaotang Du, Giwon Hong +6

Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided…