4 papers
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning
Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang +7
Large Language Models (LLMs) achieve remarkable reasoning capabilities through reinforcement learning (RL) post-training. However, existing RL post-training commonly relies on unif…
ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning
Elie Abboud, Oren Gal
Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes reward design especially delicate…
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Jonathan Steinberg, Oren Gal
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing…
Semantic Denial of Service in LLM-controlled robots
Jonathan Steinberg, Oren Gal
Safety-oriented instruction-following is supposed to keep LLM-controlled robots safe. We show it also creates an availability attack surface. By injecting short safety-plausible ph…