collaborators

5 papers

cs.LG2026

Continual Learning Mechanisms Compose for Long-Horizon Memorization

Zheyuan Zhang, Alvin Zhang, Daniel Khashabi +1

Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memoriz…

cs.LG2026

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher

Arda Uzunoglu, Alvin Zhang, Daniel Khashabi

Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view this primarily as a data sele…

cs.CL2026

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents

Zheyuan Zhang, Zehao Wen, Alvin Zhang +4

For agents to learn continuously from interaction with the world at test time, they must be able to explore effectively, acquire new world knowledge and skills, retain relevant epi…

cs.AI2026

Conformal Thinking: Risk Control for Reasoning on a Compute Budget

Xi Wang, Anushri Suresh, Alvin Zhang +6

Reasoning Large Language Models (LLMs) enable test-time scaling, with dataset-level accuracy improving as the token budget increases, motivating adaptive reasoning -- spending toke…

cs.CL2025

Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback

Dongwei Jiang, Alvin Zhang, Andrew Wang +2

Recent studies have shown LLMs possess some ability to improve their responses when given external feedback. However, it remains unclear how effectively and thoroughly these models…