2 papers
cs.LG2026
Removing Sandbagging in LLMs by Training with Weak Supervision
Emil Ryd, Henning Bartsch, Julian Stastny +2
As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more cap…
cs.CL2025
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
Sharan Maiya, Henning Bartsch, Nathan Lambert +1
The character of the "AI assistant" persona generated by modern chatbot large language models influences both surface-level behavior and apparent values, beliefs, and ethics. These…