2 papers
cs.AI2025
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
Antonio-Gabriel Chacón Menke, Phan Xuan Tan, Eiji Kamioka
Recent work has highlighted the importance of monitoring chain-of-thought reasoning for AI safety; however, current approaches that analyze textual reasoning steps can miss subtle…
cs.LG2025
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
Antonio-Gabriel Chacón Menke, Phan Xuan Tan
Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-cri…