From the 1 of 6 linked papers with an AI index.
6 papers
MESHA: Mechanism-Enforced Sequential Halving for Strategic Linear Bandits
Xin Li, Zixin Zhong
The paper proposes MESHA, a mechanism‑enforced sequential halving algorithm that identifies the best arm in linear bandits where arms may misreport their features strategically, an…
Replay What Matters: Off-Policy Replay for Efficient LLM Reinforcement Unlearning
Zirui Pang, Chenlong Zhang, Haosheng Tan +3
LLM unlearning has emerged as a cost-effective alternative to full retraining for removing hazardous knowledge from pretrained models while preserving general utility. Recent RL-ba…
Learning Multi-Agent Communication Protocol: Study on Information Entropy Efficiency in MARL
Xinren Zhang, Zixin Zhong, Jiadong Yu
Multi-Agent Systems (MAS) have emerged as a fundamental paradigm for distributed problem-solving, where autonomous agents collaborate to achieve complex objectives. Within this fra…
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Yuhang He, Haodong Wu, Siyi Liu +7
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment…
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
Yunlong Hou, Zixin Zhong, Vincent Y. F. Tan
We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured by the classic regret minimiz…
Learning Efficient Communication Protocols for Multi-Agent Reinforcement Learning
Xinren Zhang, Jiadong Yu, Zixin Zhong
Multi-Agent Systems (MAS) have emerged as a powerful paradigm for modeling complex interactions among autonomous entities in distributed environments. In Multi-Agent Reinforcement…