5 papers
SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator ove…
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattere…
Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck
Chenyu Zhou, Qiliang Jiang, Xu Zhou
Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a langua…
Certified Speculative Execution for Untrusted AI Agents
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM…
The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study the opposite signal: a deterministi…