15 papers
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5
The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…
Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
Guo Yu, Wenlin Liu, Yulan Hu +3
On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generated trajectories and dense token-l…
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
Xiang Cheng, Yulan Hu, Lulu Zheng +3
Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single user, where the field is approachin…
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
Lulu Zheng, Wenjin Yang, Xiangwen Zhang +4
Multi-stakeholder tasks require one output to satisfy users with conflicting preferences. Holistic LLM judges conflate utility estimation and utility aggregation, yielding unstable…
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs
Ziqiao Shang, Lingyue Ge, Ling-Yue Ge +12
Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for r…
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen +4
On-policy self-distillation (self-OPD) densifies reinforcement learning with verifiable rewards (RLVR) by letting a policy teach itself under privileged context. We find that when…