works on

From the 1 of 15 linked papers with an AI index.

collaborators

15 papers

cs.AI2026

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

ZhiYan Hou, Xinyu Tang, Hongyan An +9

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals…

cs.LG2026

Continual Learning in Transition

Zhiyan Hou, Dan Zhang, Tao Feng +11

Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architect…

cs.CV2026

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

Shijie Wang, Xiangzhao Hao, Yueti Li +3

Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings through single forward encoding…

cs.CL2026

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Xinyu Tang, Qianggang Cao, Yurou Liu +13

The paper introduces a training pipeline that scales zero‑reinforcement‑learning to a trillion‑parameter language model, revealing emergent chain‑of‑thought reasoning abilities and…

cs.CL2026

GraphPO: Graph-based Policy Optimization for Reasoning Models

Yuliang Zhan, Xinyu Tang, Jian Li +7

Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses indepe…

cs.AI2026

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

Xiaochong Lan, Pu Ning, Quan Chen +8

Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inhe…