works on

From the 1 of 24 linked papers with an AI index.

collaborators

24 papers

cs.CL2026

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Xinyu Geng, Xuanhua He, Sixiang Chen +7

The paper introduces DeepSearch-World, a deterministic, verifiable web environment, and DeepSearch-Evolve, a self‑distillation framework that lets web search agents improve from th…

cs.AI2026

Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning

Rui Liu, Dian Yu, Tong Zheng +8

Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inpu…

cs.LG2026

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents

Yujun Zhou, Kehan Guo, Haomin Zhuang +8

Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated…

cs.AI2026

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning

Jiangnan Xia, Yucheng Shi, Yu Yang +3

Reinforcement learning has become a key paradigm for eliciting reasoning abilities in large language models, where exploration is crucial for discovering effective solution traject…

cs.CL2026

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization

Cihan Xiao, Yiwen Shao, Chenxing Li +5

Audio and omni-modal large language models exhibit impressive cross-modal reasoning capabilities. However, applying standard reinforcement learning post-training algorithms to thes…

cs.LG2026

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

Huilin Zhou, Jian Zhao, Yilu Zhong +7

Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static…