3 papers
cs.CL2026
Local Edits, Global Ripples: Replay-Informed Policy Adaptation for Workflow Synthesis
Manqing Mao, Hong Wang, Samson Koelle +8
Prompt-policy editing offers a practical way to improve agents that synthesize executable workflows without updating the underlying model. However, persistent prompt editing has tw…
cs.LG2026
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning
Zheyuan Zhang, Manqing Mao, Hong Wang +8
Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of ro…
cs.CL2024
Multi-User Chat Assistant (MUCA): a Framework Using LLMs to Facilitate Group Conversations
Manqing Mao, Paishun Ting, Yijian Xiang +3
Recent advancements in large language models (LLMs) have provided a new avenue for chatbot development. Most existing research, however, has primarily centered on single-user chatb…