5 papers
Coherence, charity and triangulation in statistical modelling
David J. T. Sumpter
Bayesian statistics rests on a few familiar distinctions: frequentist vs. Bayesian, objective versus subjective probability, a model versus the data it is fitted to, a prior versus…
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
John T. Halloran, Noopur S. Bhatt
Large language models (LLMs) are highly susceptible to backdoor attacks (BAs), wherein training samples are poisoned using trigger-based harmful content. Furthermore, existing defe…
Leveraging RAG for Training-Free Alignment of LLMs
John T. Halloran
Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to enable safety guardrails and alig…
Understanding the Effects of Safety Unalignment on Large Language Models
John T. Halloran
Safety alignment has become a critical step to ensure LLMs refuse harmful requests while providing helpful and harmless responses. However, despite the ubiquity of safety alignment…
Mamba State-Space Models Are Lyapunov-Stable Learners
John T. Halloran, Manbir Gulati, Paul F. Roysdon
Mamba state-space models (SSMs) have recently outperformed state-of-the-art (SOTA) Transformer large language models (LLMs) in various tasks and been widely adapted. However, a maj…