Showing 2026Show all
2 papers · 1 filter
cs.LG2026
medR: Reward Engineering for Clinical Offline Reinforcement Learning via Tri-Drive Potential Functions
Qianyi Xu, Gousia Habib, Feng Wu +5
Reinforcement Learning (RL) offers a powerful framework for optimizing dynamic treatment regimes (DTRs). However, clinical RL is fundamentally bottlenecked by reward engineering: t…
cs.CL2026
S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
Yanrui Du, Sendong Zhao, Yibo Gao +9
Large language models (LLMs) equipped with chain-of-thought (CoT) achieve strong performance and offer a window into LLM behavior. However, recent evidence suggests that improvemen…