advantage shaping 1contrastive learning 1on-policy distillation 1policy optimization 1token-level correctness 1
From the 1 of 6 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
Song Wang, Zihan Chen, Peng Wang +5
Retrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or special…
cs.CL2025
The End of Manual Decoding: Towards Truly End-to-End Language Models
Zhichao Wang, Dongyang Ma, Xinting Huang +6
The "end-to-end" label for LLMs is a misnomer. In practice, they depend on a non-differentiable decoding process that requires laborious, hand-tuning of hyperparameters like temper…
cs.CL2024
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
Deng Cai, Huayang Li, Tingchen Fu +11
Despite the general capabilities of pre-trained large language models (LLMs), they still need further adaptation to better serve practical applications. In this paper, we demonstra…