collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Rubric-to-Code Credit Assignment for Reinforcement Learning

Rui Jin, Jikai Chen, Yihan Chen +6

Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation,…

cs.AI2026

MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG

Qiuyi Qi, Tian Liang, Jiamu Wang +7

Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching and when to answer. Existing RL-based methods rely on external supervisio…

cs.AI2026

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

Zechuan Wang, Siyuan Lu, Hongxuan Zhang +3

Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignm…

cs.AI2026

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

Qiuyi Qi, Tian Liang, Mutian Bao +8

Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to traject…

cs.AI2026

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

Qiuyi Qi, Jinjian Zhang, Mutian Bao +9

Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their r…