collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

ProofVerifier: A Scalable, Diversity-Driven Framework for Natural-Language Proof Verification

Haotong Yang, Zitong Wang, Shijia Kang +7

While large language models (LLMs) have achieved strong performance on mathematical problems with verifiable answers, many advanced problems are proof-based and require evaluating…

cs.CL2026

SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding

Shuyang Hou, Yi Hu, Muhan Zhang

Recent advancements in large language models (LLMs) have significantly enhanced their reasoning capabilities. However, they continue to struggle with basic character-level tasks, s…

cs.CL2026

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

Yi Hu, Jiaqi Gu, Ruxin Wang +6

Reinforcement learning (RL) has catalyzed the emergence of Large Reasoning Models (LRMs) that have pushed reasoning capabilities to new heights. While their performance has garnere…

cs.CL2025

What Affects the Effective Depth of Large Language Models?

Yi Hu, Cai Zhou, Muhan Zhang

The scaling of large language models (LLMs) emphasizes increasing depth, yet performance gains diminish with added layers. Prior work introduces the concept of "effective depth", a…

cs.CL2025

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Shi Qiu, Shaoyang Guo, Zhuo-Yang Song +51

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed e…