works on

From the 1 of 14 linked papers with an AI index.

most citedRethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 7 of their papers we have counts for

collaborators

14 papers

cs.AI2026

Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping

Yiwen Zhang, Xiaodong Yan, Zhenyu Huang +6

Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the code generation capabilities of Large Language Models (LLMs). However, th…

cs.CL2026

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Xinyu Tang, Qianggang Cao, Yurou Liu +13

The paper introduces a training pipeline that scales zero‑reinforcement‑learning to a trillion‑parameter language model, revealing emergent chain‑of‑thought reasoning abilities and…

cs.CL2026

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

Ang Li, Ben Liu, Bin Han +215

Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve,…

cs.AI2026

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

Xiaochong Lan, Pu Ning, Quan Chen +8

Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inhe…

cs.AI2026

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Xinyu Lu, Tianshu Wang, Pengbo Wang +8

Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next-level capability: whether m…

cs.CL2026

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

Yuchen Yan, Liang Jiang, Jin Jiang +7

Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, context length limits, and degraded rea…