3 papers
cs.LG2025
A Provable Approach for End-to-End Safe Reinforcement Learning
Akifumi Wachi, Kohei Miyaguchi, Takumi Tanabe +2
A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing sa…
cs.AI2025
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
Thien Q. Tran, Akifumi Wachi, Rei Sato +2
Safety alignment is an essential research topic for real-world AI applications. Despite the multifaceted nature of safety and trustworthiness in AI, current safety alignment method…
cs.LG2024
Stepwise Alignment for Constrained Language Model Policy Optimization
Akifumi Wachi, Thien Q. Tran, Rei Sato +2
Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment…