3 papers
cs.CL2026
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
Zheyuan Zhang, Zehao Wen, Alvin Zhang +4
For agents to learn continuously from interaction with the world at test time, they must be able to explore effectively, acquire new world knowledge and skills, retain relevant epi…
cs.CL2025
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
Dongwei Jiang, Alvin Zhang, Andrew Wang +2
Recent studies have shown LLMs possess some ability to improve their responses when given external feedback. However, it remains unclear how effectively and thoroughly these models…
cs.AI2025
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
Dongwei Jiang, Guoxuan Wang, Yining Lu +5
The reasoning steps generated by LLMs might be incomplete, as they mimic logical leaps common in everyday communication found in their pre-training data: underlying rationales are…