2 papers
cs.AI2026
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
Sheng Ren, Yadong Wang, Naiqiang Tan +7
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when…
cs.LG2026
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
Qin-Wen Luo, Sheng Ren, Xiang Chen +4
Chain-of-Thought (CoT) has substantially empowered Large Language Models (LLMs) to tackle complex reasoning tasks, yet the verbose nature of explicit reasoning steps incurs prohibi…