3 papers
cs.LG2026
ReNCE: Learning to Reason by Noise Contrastive Estimation
Wenzheng Zhang, Karl Stratos
GRPO is a standard approach to endowing pretrained LLMs with reasoning capabilities. It estimates the advantage of an outcome from a group of outcomes, and promotes those with…
cs.CL2025
ImpRAG: Retrieval-Augmented Generation with Implicit Queries
Wenzheng Zhang, Xi Victoria Lin, Karl Stratos +2
Retrieval-Augmented Generation (RAG) systems traditionally treat retrieval and generation as separate processes, requiring explicit textual queries to connect them. This separation…
cs.CL2025
Rethinking Reflection in Pre-Training
Essential AI, :, Darsh J Shah +26
A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develop…