2 papers
cs.CL2026
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Maciej ChrabÄ szcz, Aleksander Szymczyk, Marcin Sendera +2
Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's f…
cs.CV2026
Conditioned Activation Transport for T2I Safety Steering
Maciej ChrabÄ szcz, Aleksander Szymczyk, Jan DubiÅski +3
Despite their impressive capabilities, current Text-to-Image (T2I) models remain prone to generating unsafe and toxic content. While activation steering offers a promising inferenc…