5 papers · 1 filter
Symbolic and Abstractive Reasoning with Complex Visual Queries
Yichi Zhang, Jingdian Lu, Zhuo Chen +4
Understanding and reasoning over abstract visual content remains a challenge for current multi-modal large language models (MLLMs). In this paper, we explore a novel abstract data…
MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment
Zihao Liu, Hantao Zhou, Jiguo Li +5
User simulators are essential for the scalable training and evaluation of interactive AI systems. However, existing approaches often rely on shallow user profiling, struggle to mai…
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
Nonghai Zhang, Weitao Ma, Zhanyu Ma +5
Group Relative Policy Optimization (GRPO) significantly enhances the reasoning performance of Large Language Models (LLMs). However, this success heavily relies on expensive extern…
UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning
Feng Zhang, Shijia Li, Chunmao Zhang +7
User simulators serve as the critical interactive environment for agent post-training, and an ideal user simulator generalizes across domains and proactively engages in negotiation…
Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management
Weitao Ma, Xiaocheng Feng, Lei Huang +7
Effective memory management is essential for large language model agents to navigate long-horizon tasks. Recent research has explored using Reinforcement Learning to develop specia…