37 papers
MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques
Jiabao Zhuang, Changhao Jiang, Hanchen Wang +11
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligni…
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Junjie Ye, Zhuohui Sheng, Shaofan Liu +12
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps)…
IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
Dingwei Zhu, Jiahan Li, Chengjun Pan +22
Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history sca…
HSD: Hybrid Hindsight Self-Distillation
Qiye Cai, Yichuan Ma, Linyang Li +7
Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level…
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
Zhiheng Xi, Dingwen Yang, Jiaqi Liu +21
Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic eval…
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
Honglin Guo, Qi Zhang, Yu Zhang +6
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating spa…