43 citations · 55 across the 35 of their papers we have counts for
39 papers
Continual Learning Mechanisms Compose for Long-Horizon Memorization
Zheyuan Zhang, Alvin Zhang, Daniel Khashabi +1
Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memoriz…
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning
Zheyuan Zhang, Manqing Mao, Hong Wang +8
Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of ro…
HoosierHelp: Benchmarking LLM Agents for Social Service Navigation
Yiyang Li, Weixiang Sun, Tianyi Ma +3
Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interfa…
Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier
Kaiwen Shi, Zheyuan Zhang, Han Bao +2
Modern agent systems can turn uncertainty into overconfidence. Fragile upstream decisions are often exposed to downstream components as clean intermediate artifacts, while the unce…
SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment
Kaiwen Shi, Zheyuan Zhang, Yanfang Ye
Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's sampled behavior. We study verba…
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
Zheyuan Zhang, Kaiwen Shi, Han Bao +3
Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from model-generated outputs but l…