1 citations · 1 across the 7 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025★ 1 cited
SSRL: Self-Search Reinforcement Learning
Yuchen Fan, Kaiyan Zhang, Heng Zhou +15
We investigate the potential of large language models (LLMs) to serve as efficient simulators for agentic search tasks in reinforcement learning (RL), thereby reducing dependence o…
cs.CL2025
Through the Valley: Path to Effective Long CoT Training for Small Language Models
Renjie Luo, Jiaxi Li, Chen Huang +1
Long chain-of-thought (CoT) supervision has become a common strategy to enhance reasoning in language models. While effective for large models, we identify a phenomenon we call Lon…
cs.CL2025
Delta -- Contrastive Decoding Mitigates Text Hallucinations in Large Language Models
Cheng Peng Huang, Hao-Yuan Chen
Large language models (LLMs) demonstrate strong capabilities in natural language processing but remain prone to hallucinations, generating factually incorrect or fabricated content…