collaborators

5 papers

cs.CL2026

SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding

Shuyang Hou, Yi Hu, Muhan Zhang

Recent advancements in large language models (LLMs) have significantly enhanced their reasoning capabilities. However, they continue to struggle with basic character-level tasks, s…

cs.CL2026

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

Yi Hu, Jiaqi Gu, Ruxin Wang +6

Reinforcement learning (RL) has catalyzed the emergence of Large Reasoning Models (LRMs) that have pushed reasoning capabilities to new heights. While their performance has garnere…

cs.CL2025

What Affects the Effective Depth of Large Language Models?

Yi Hu, Cai Zhou, Muhan Zhang

The scaling of large language models (LLMs) emphasizes increasing depth, yet performance gains diminish with added layers. Prior work introduces the concept of "effective depth", a…

cs.LG2025

The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling

Shijia Kang, Muhan Zhang

Reinforcement learning (RL) has been pivotal in enhancing the reasoning capabilities of large language models (LLMs), but it often suffers from limited exploration and entropy coll…

cs.CL2025

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Shi Qiu, Shaoyang Guo, Zhuo-Yang Song +51

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed e…