benchmark evaluation 1intermediate visual states 1large language models 1multimodal models 1visual reasoning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
Haoze Liu, Run Liu, Haiying Xu +6
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM per…
cs.CV2026
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Siyu Yan, Zhuoran Yan, Haiying Xu +10
The paper presents See2Think, an evaluation framework and benchmark for testing whether multimodal large language models actually use intermediate visual states during reasoning, a…
cs.CV2026
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
Haiying Xu, Zihan Wang, Song Dai +3
Despite recent advances in multimodal reasoning, representing auxiliary geometric constructions remains a fundamental challenge for multimodal large language models (MLLMs). Such c…