Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
Julian Schulz
As AI systems approach dangerous capability levels where inability safety cases become insufficient, we need alternative approaches to ensure safety. This paper presents a roadmap…
cs.LG2025
Automated Feature Labeling with Token-Space Gradient Descent
Julian Schulz, Seamus Fallows
We present a novel approach to feature labeling using gradient descent in token-space. While existing methods typically use language models to generate hypotheses about feature mea…