7 papers
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
Marek Jeliński, Jan Dubiński, Maciej Chrabaszcz +1
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated t…
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Maciej Chrabąszcz, Aleksander Szymczyk, Marcin Sendera +2
Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's f…
Conditioned Activation Transport for T2I Safety Steering
Maciej Chrabąszcz, Aleksander Szymczyk, Jan Dubiński +3
Despite their impressive capabilities, current Text-to-Image (T2I) models remain prone to generating unsafe and toxic content. While activation steering offers a promising inferenc…
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation
Szymon Płotka, Gizem Mert, Maciej Chrabaszcz +2
In recent years, artificial intelligence has significantly advanced medical image segmentation. Nonetheless, challenges remain, including efficient 3D medical image processing acro…
Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
Maciej Chrabąszcz, Katarzyna Lorenc, Karolina Seweryn
Large language models (LLMs) have demonstrated impressive capabilities across various natural language processing (NLP) tasks in recent years. However, their susceptibility to jail…
Efficient LLM Moderation with Multi-Layer Latent Prototypes
Maciej Chrabąszcz, Filip Szatkowski, Bartosz Wójcik +3
Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suff…