works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.LG2026

MESHA: Mechanism-Enforced Sequential Halving for Strategic Linear Bandits

Xin Li, Zixin Zhong

The paper proposes MESHA, a mechanism‑enforced sequential halving algorithm that identifies the best arm in linear bandits where arms may misreport their features strategically, an…

cs.CL2026

Replay What Matters: Off-Policy Replay for Efficient LLM Reinforcement Unlearning

Zirui Pang, Chenlong Zhang, Haosheng Tan +3

LLM unlearning has emerged as a cost-effective alternative to full retraining for removing hazardous knowledge from pretrained models while preserving general utility. Recent RL-ba…

cs.MA2026

Learning Multi-Agent Communication Protocol: Study on Information Entropy Efficiency in MARL

Xinren Zhang, Zixin Zhong, Jiadong Yu

Multi-Agent Systems (MAS) have emerged as a fundamental paradigm for distributed problem-solving, where autonomous agents collaborate to achieve complex objectives. Within this fra…

cs.LG2026

Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR

Yuhang He, Haodong Wu, Siyi Liu +7

Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment…

cs.LG2026

On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits

Yunlong Hou, Zixin Zhong, Vincent Y. F. Tan

We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured by the classic regret minimiz…

cs.MA2025

Learning Efficient Communication Protocols for Multi-Agent Reinforcement Learning

Xinren Zhang, Jiadong Yu, Zixin Zhong

Multi-Agent Systems (MAS) have emerged as a powerful paradigm for modeling complex interactions among autonomous entities in distributed environments. In Multi-Agent Reinforcement…