activity
20242026
collaborators

10 papers

cs.CL2026

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

Kang Peng, Zhiwei Zhang, Yichen Zhang +7

Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following proc…

cs.AI2026

Expectation Alignment of Language Models for Real-World User Expectations

Miaomiao Li, Yang Wang, Bin Liang +3

Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing…

cs.LG2026

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

Zhiwei Zhang, Fei Zhao, Rui Wang +6

Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models often fall into repetitive invalid…

cs.CL2025

ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning

Yiming Du, Yifan Xiang, Bin Liang +3

Fine-tuning multi-turn dialogue systems requires high-quality supervision but often suffers from degraded performance when exposed to low-quality data. Supervision errors in early…

cs.CL2025

MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents

Yiming Du, Bingbing Wang, Yang He +7

Modern task-oriented dialogue (TOD) systems increasingly rely on large language model (LLM) agents, leveraging Retrieval-Augmented Generation (RAG) and long-context capabilities fo…

cs.MM2025

Reply with Sticker: New Dataset and Model for Sticker Retrieval

Bin Liang, Bingbing Wang, Zhixin Bai +6

Using stickers in online chatting is very prevalent on social media platforms, where the stickers used in the conversation can express someone's intention/emotion/attitude in a viv…