4 papers
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
Yibo Yan, Shen Wang, Jiahao Huo +13
As the field of Multimodal Large Language Models (MLLMs) continues to evolve, their potential to revolutionize artificial intelligence is particularly promising, especially in addr…
LLM Agents for Education: Advances and Applications
Zhendong Chu, Shen Wang, Jian Xie +8
Large Language Model (LLM) agents are transforming education by automating complex pedagogical tasks and enhancing both teaching and learning processes. In this survey, we present…
ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
Jian Xie, Zhendong Chu, Aoxiao Zhong +5
Large Reasoning Models (LRMs) often suffer from the ``over-thinking'' problem, generating unnecessarily long reasoning on simple tasks. Some strategies have been proposed to mitiga…
League: Leaderboard Generation on Demand
Jian Wu, Jiayu Zhang, Dongyuan Li +5
This paper introduces Leaderboard Auto Generation (LAG), a novel and well-organized framework for automatic generation of leaderboards on a given research topic in rapidly evolving…