Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Shelly Bensal, Umar Jamil, Christopher Bryant +5
We explore a method for improving the performance of large language models through self-reflection and reinforcement learning. By incentivizing the model to generate better self-re…
cs.CL2024
Writing in the Margins: Better Inference Pattern for Long Context Retrieval
Melisa Russak, Umar Jamil, Christopher Bryant +4
In this paper, we introduce Writing in the Margins (WiM), a new inference pattern for Large Language Models designed to optimize the handling of long input sequences in retrieval-o…