1 citations · 1 across the 14 of their papers we have counts for
Showing 2026Show all
3 papers · 1 filter
cs.LG2026
On the Complexity of Offline Reinforcement Learning with -Approximation and Partial Coverage
Haolin Liu, Braham Snyder, Chen-Yu Wei
We study offline reinforcement learning under -approximation and partial coverage, a setting that motivates practical algorithms such as Conservative -Learning (CQL; Ku…
cs.LG2026
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
Haolin Liu, Dian Yu, Sidi Lu +6
Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…
cs.CL2026
RelayLLM: Efficient Reasoning via Collaborative Decoding
Chengsong Huang, Tong Zheng, Langlin Huang +3
Large Language Models (LLMs) for complex reasoning is often hindered by high computational costs and latency, while resource-efficient Small Language Models (SLMs) typically lack t…