1 paper · 1 filter
Antonio-Gabriel Chacón Menke, Phan Xuan Tan, Eiji Kamioka
Recent work has highlighted the importance of monitoring chain-of-thought reasoning for AI safety; however, current approaches that analyze textual reasoning steps can miss subtle…