activity
20242026
collaborators

54 papers

cs.CL2026

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Junjie Ye, Zhuohui Sheng, Shaofan Liu +12

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps)…

cs.AI2026

MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

Jiajun Dong, Yutao Hu, Fengrui Fan +3

Large language model (LLM) agents must retain and use cross-step information to act coherently in long-horizon tasks. Existing methods improve memory accessibility, yet action-rele…

cs.CL2026

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

Dingwei Zhu, Jiahan Li, Chengjun Pan +22

Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history sca…

cs.LG2026

VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training

Dingwei Zhu, Shihan Dou, Zhiheng Xi +16

Reinforcement Learning (RL) in real-world environments often suffers from ambiguous or incomplete reward supervision, which undermines policy stability and generalization. Such noi…

cs.CL2026

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

Changze Lv, Jie Zhou, Wentao Zhao +12

Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward signals. Accordingly, rubric-bas…

cs.CL2026

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

Yujiong Shen, Yajie Yang, Zhiheng Xi +17

Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely overlook agents' ability to orches…