3 papers
cs.CR2026
Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study
Paria Mehrbod, Boris Knyazev, Guy Wolf +2
Despite extensive safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attribute…
cs.LG2025
Test Time Adaptation Using Adaptive Quantile Recalibration
Paria Mehrbod, Pedro Vianna, Geraldin Nanfack +2
Domain adaptation is a key strategy for enhancing the generalizability of deep learning models in real-world scenarios, where test distributions often diverge significantly from th…
cs.LG2025
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
Vaibhav Singh, Paul Janson, Paria Mehrbod +4
The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has em…