3 papers
cs.AI2026
Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
Aochong Oliver Li, Tanya Goyal
Reasoning LLMs are trained to verbalize their reasoning process, yielding strong gains on complex tasks. This transparency also opens a promising direction: multiple reasoners can…
cs.LG2026
Better LLM Reasoning via Dual-Play
Zhengxin Zhang, Chengyu Huang, Aochong Oliver Li +1
Large Language Models (LLMs) have achieved remarkable progress through Reinforcement Learning with Verifiable Rewards (RLVR), yet still rely heavily on external supervision (e.g.,…
cs.CL2025
Memorization vs. Reasoning: Updating LLMs with New Knowledge
Aochong Oliver Li, Tanya Goyal
Large language models (LLMs) encode vast amounts of pre-trained knowledge in their parameters, but updating them as real-world information evolves remains a challenge. Existing met…