collaborators

6 papers

cs.AI2026

To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models

Haoqing Wang, Xiang Long, Ziheng Li +3

Reinforcement Learning with Verifiable Rewards (RLVR) plays a key role in stimulating the explicit reasoning capability of Large Language Models (LLMs). We can achieve expert-level…

cs.CL2026

Self-Manager: Parallel Agent Loop for Long-form Deep Research

Yilong Xu, Zhi Zheng, Xiang Long +2

Long-form deep research requires multi-faceted investigations over extended horizons to get a comprehensive report. When handling such complex tasks, existing agents manage context…

cs.CL2025

An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs

Linyue Ma, Yilong Xu, Xiang Long +1

Search augmentation empowers Large Language Models with retrieval capabilities to overcome the limitations imposed by static parameters. Recently, Reinforcement Learning leverages…

cs.LG2025

APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation

Yuzhen Zhou, Jiajun Li, Yusheng Su +15

Reinforcement learning (RL) has become a cornerstone in advancing large-scale pre-trained language models (LLMs). Successive generations, including GPT-o series, DeepSeek-R1, Kimi-…

cs.CL2025

MiniCPM4: Ultra-Efficient LLMs on End Devices

MiniCPM Team, Chaojun Xiao, Yuxuan Li +80

This paper introduces MiniCPM4, a highly efficient large language model (LLM) designed explicitly for end-side devices. We achieve this efficiency through systematic innovation in…

cs.CL2025

RAVine: Reality-Aligned Evaluation for Agentic Search

Yilong Xu, Xiang Long, Zhi Zheng +1

Agentic search, as a more autonomous and adaptive paradigm of retrieval augmentation, is driving the evolution of intelligent search systems. However, existing evaluation framework…