3 papers
cs.LG2026
Efficient Refusal Ablation in LLM through Optimal Transport
Geraldin Nanfack, Eugene Belilovsky, Elvis Dohmatob
Safety-aligned language models refuse harmful requests through learned refusal behaviors encoded in their internal representations. Recent activation-based jailbreaking methods cir…
cs.LG2025
Test Time Adaptation Using Adaptive Quantile Recalibration
Paria Mehrbod, Pedro Vianna, Geraldin Nanfack +2
Domain adaptation is a key strategy for enhancing the generalizability of deep learning models in real-world scenarios, where test distributions often diverge significantly from th…
cs.LG2025
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
Geraldin Nanfack, Eugene Belilovsky
Deep learning models frequently exploit spurious features in training data to achieve low training error, often resulting in poor generalization when faced with shifted testing dis…