9 papers
Context-Aware RL for Agentic and Multimodal LLMs
Peiyang Xu, Bangzheng Li, Sijia Liu +4
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool…
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
Zhaonan Li, Kyle R. Chickering, Bangzheng Li +13
A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties…
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
Kyle R. Chickering, Bangzheng Li, Muhao Chen
Multimodal Large Language Models (MLLMs) encode images into visual tokens, aligning visual and textual signals within a shared latent space to facilitate crossmodal representation…
Reinforced Attention Learning
Bangzheng Li, Jianmo Ni, Chen Qu +5
Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multi…
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
Rui Cai, Bangzheng Li, Xiaofei Wen +2
Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet often exhibit poor robustness when exposed to spurious modality interference, such as…
Unbiased Visual Reasoning with Controlled Visual Inputs
Zhaonan Li, Shijie Lu, Fei Wang +11
End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone whe…