Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
Hongwei Chen, Yishu Lei, Dan Zhang +10
Test-time scaling has emerged as a promising paradigm in language modeling, wherein additional computational resources are allocated during inference to enhance model performance.…
cs.CL2024
Tool-Augmented Reward Modeling
Lei Li, Yekun Chai, Shuohuan Wang +4
Reward modeling (a.k.a., preference modeling) is instrumental for aligning large language models with human preferences, particularly within the context of reinforcement learning f…