5 papers
SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding
Shuyang Hou, Yi Hu, Muhan Zhang
Recent advancements in large language models (LLMs) have significantly enhanced their reasoning capabilities. However, they continue to struggle with basic character-level tasks, s…
Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures
Yi Hu, Jiaqi Gu, Ruxin Wang +6
Reinforcement learning (RL) has catalyzed the emergence of Large Reasoning Models (LRMs) that have pushed reasoning capabilities to new heights. While their performance has garnere…
What Affects the Effective Depth of Large Language Models?
Yi Hu, Cai Zhou, Muhan Zhang
The scaling of large language models (LLMs) emphasizes increasing depth, yet performance gains diminish with added layers. Prior work introduces the concept of "effective depth", a…
The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling
Shijia Kang, Muhan Zhang
Reinforcement learning (RL) has been pivotal in enhancing the reasoning capabilities of large language models (LLMs), but it often suffers from limited exploration and entropy coll…
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Shi Qiu, Shaoyang Guo, Zhuo-Yang Song +51
Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed e…