collaborators

10 papers

cs.AI2026

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

Chenyu Zhou, Qiliang Jiang, Shuning Wu +1

A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator ove…

cs.LG2026

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

Chenyu Zhou, Qiliang Jiang, Shuning Wu +1

Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattere…

cs.CV2026

Pixel Decodability Is Not a Compression Signal: Causally Evaluating Importance Proxies for Visual KV-Cache Eviction

Chenyu Zhou, Qiliang Jiang, Shuning Wu +1

Vision-language models retain a substantial amount of pixel-decodable visual content in their visual key-value cache. We show, in our setting, that this retention is task-inert: ac…

cs.CR2026

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

Yitian Zhou, Jingyu Zheng, Qiliang Jiang +6

Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents…

cs.CL2026

Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

Chenyu Zhou, Qiliang Jiang, Xu Zhou

Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a langua…

cs.CR2026

Certified Speculative Execution for Untrusted AI Agents

Chenyu Zhou, Qiliang Jiang, Shuning Wu +1

Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM…