collaborators

6 papers

cs.CL2026

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

Kang Peng, Zhiwei Zhang, Yichen Zhang +7

Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following proc…

cs.AI2026

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows

Hongliang Li, Yijin Liu, Zhiwei Zhang +5

While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across ex…

cs.HC2026

PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications

Yifan Simon Liu, Qianfeng Wen, Yilan Fan +40

Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and…

cs.AI2026

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Xiaomin Li, Yuexing Hao, Jianheng Hou +90

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…

cs.AI2026

Expectation Alignment of Language Models for Real-World User Expectations

Miaomiao Li, Yang Wang, Bin Liang +3

Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing…

cs.LG2026

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

Zhiwei Zhang, Fei Zhao, Rui Wang +6

Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models often fall into repetitive invalid…