Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
Zhengyang Zhao, Lu Ma, Wentao Zhang
Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying model remain unchanged by the…
cs.CL2026
Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning
Lexiang Tang, Weihao Gao, Bingchen Zhao +4
Recent work on test-time scaling for large language model (LLM) reasoning typically assumes that allocating more inference-time computation uniformly improves correctness. However,…
cs.CL2025
DARO: Difficulty-Aware Reweighting Policy Optimization
Jingyu Zhou, Lu Ma, Hao Liang +3
Recent advances in large language models (LLMs) have shown that reasoning ability can be significantly enhanced through Reinforcement Learning with Verifiable Rewards (RLVR). Group…