Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
TRE: Encouraging Exploration in the Trust Region
Chao Huang, Yujing Lu, Quangang Li +8
Entropy regularization is a standard technique in reinforcement learning (RL) to enhance exploration, yet it yields negligible effects or even degrades performance in Large Languag…
cs.CL2025
Electronic Circuit Principles of Large Language Models
Qiguang Chen, Libo Qin, Jinhao Liu +6
Large language models (LLMs) such as DeepSeek-R1 have achieved remarkable performance across diverse reasoning tasks. To uncover the principles that govern their behaviour, we intr…
cs.CL2024
Exploring Hybrid Question Answering via Program-based Prompting
Qi Shi, Han Cui, Haofeng Wang +3
Question answering over heterogeneous data requires reasoning over diverse sources of data, which is challenging due to the large scale of information and organic coupling of heter…