4 papers
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling
Jiarui Zhao, Rongzhi Zhang, Lingchuan Liu +3
Search agent benchmarks exemplified by BrowseComp have rapidly saturated over the past year, with the strongest models surpassing 90% accuracy. Since these benchmarks are predomina…
Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
Ruichao Mao, Zhou Fang, Teng Guo +9
User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large lan…
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
Haoming Yang, Ke Ma, Xiaojun Jia +3
Despite the remarkable performance of Large Language Models (LLMs), they remain vulnerable to jailbreak attacks, which can compromise their safety mechanisms. Existing studies ofte…
Learning from Peers in Reasoning Models
Tongxu Luo, Wenyu Du, Jiaxi Bi +5
Large Reasoning Models (LRMs) have the ability to self-correct even when they make mistakes in their reasoning paths. However, our study reveals that when the reasoning process sta…