6 papers
Safety Hacking in Constrained Best-of- Inference-time Scaling
Akifumi Wachi, Takumi Tanabe, Youhei Akimoto
Inference-time pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with the highest learned reward. We show that…
Evolutionary Bilevel Reward Shaping for Generalization in Reinforcement Learning
Ekasit Usaratniwart, Xilin Gao, Marc Ong +1
Reinforcement learning (RL) often suffers from performance degradation when deployed in environments that differ from those encountered during training. Existing techniques such as…
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
Mikoto Kudo, Takumi Tanabe, Akifumi Wachi +1
Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcement learning (RL), where a leader agent…
Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
Shigeki Kusaka, Keita Saito, Mikoto Kudo +3
Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO a…
A Provable Approach for End-to-End Safe Reinforcement Learning
Akifumi Wachi, Kohei Miyaguchi, Takumi Tanabe +2
A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing sa…
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
Thien Q. Tran, Akifumi Wachi, Rei Sato +2
Safety alignment is an essential research topic for real-world AI applications. Despite the multifaceted nature of safety and trustworthiness in AI, current safety alignment method…