2 papers
cs.AI2026
Geometry-Guided Constraint Learning for LLM Safety Classification
Fumiaki Uehara, Koo Imai, Masato Tsutsumi +3
Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show that sparse autoencoder (SAE)…
cs.LG2026
Generalization Limits of Reinforcement Learning Alignment
Haruhi Shida, Koo Imai, Keigo Kansa
The safety of large language models (LLMs) relies on alignment techniques such as reinforcement learning from human feedback (RLHF). However, recent theoretical analyses suggest th…