Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Entropy After </Think> for reasoning model early exiting
Xi Wang, James McInerney, Lequn Wang +1
Reasoning LLMs show improved performance with longer chains of thought. However, recent work has highlighted their tendency to overthink, continuing to revise answers even after re…
cs.LG2023
Off-Policy Evaluation for Large Action Spaces via Policy Convolution
Noveen Sachdeva, Lequn Wang, Dawen Liang +2
Developing accurate off-policy estimators is crucial for both evaluating and optimizing for new policies. The main challenge in off-policy estimation is the distribution shift betw…