collaborators

10 papers

cs.AI2026

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection

Chaoran Chen, Vy Nguyen, Ziji Zhang +7

Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robu…

cs.CL2026

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

Yimeng Zhang, Yingying Zhuang, Ziyi Wang +12

Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However,…

cs.CL2026

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8

Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practic…

cs.SE2026

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

Yuxuan Lu, Ziyi Wang, Yingzhou Lu +12

Training tool-calling agents requires large-scale trajectory data with verifiable labels, yet existing approaches either synthesize environments that diverge from real API behavior…

cs.CL2026

OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

Ziyi Wang, Yuxuan Lu, Wenbo Li +13

Can large language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generating ``believable'' human behavio…

cs.IR2026

LLM Agents Enable User-Governed Personalization Beyond Platform Boundaries

Jiacheng Lin, Kun Qian, Arvind Srinivasan +15

Personalization today is fundamentally platform-centric: services build user representations from the behavioral fragments they observe. Yet no platform can construct a complete pi…