2 papers
cs.CL2026
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Maciej Chrabąszcz, Aleksander Szymczyk, Marcin Sendera +2
Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's f…
cs.CV2026
Conditioned Activation Transport for T2I Safety Steering
Maciej Chrabąszcz, Aleksander Szymczyk, Jan Dubiński +3
Despite their impressive capabilities, current Text-to-Image (T2I) models remain prone to generating unsafe and toxic content. While activation steering offers a promising inferenc…