collaborators

7 papers

cs.CL2026

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

Yimeng Zhang, Yingying Zhuang, Ziyi Wang +12

Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However,…

cs.CL2026

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8

Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practic…

cs.SE2026

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

Yuxuan Lu, Ziyi Wang, Yingzhou Lu +12

Training tool-calling agents requires large-scale trajectory data with verifiable labels, yet existing approaches either synthesize environments that diverge from real API behavior…

cs.CL2026

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

Ziyi Wang, Yuxuan Lu, Yimeng Zhang +12

Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized settings with general, fixed, and…

cs.CL2026

Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning

Yimeng Zhang, Tian Wang, Jiri Gesi +14

Large Language Models (LLMs) have recently demonstrated strong potential in generating 'believable human-like' behavior in web environments. Prior work has explored augmenting trai…

cs.CY2025

See, Think, Act: Online Shopper Behavior Simulation with VLM Agents

Yimeng Zhang, Jiri Gesi, Ran Xue +10

LLMs have recently demonstrated strong potential in simulating online shopper behavior. Prior work has improved action prediction by applying SFT on action traces with LLM-generate…