3 papers
cs.LG2026
Predicting LLM Safety Before Release by Simulating Deployment
Marcus Williams, Hannah Sheahan, Cameron Raymond +8
Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model beha…
cs.AI2026
Context Attribution with Multi-Armed Bandit Optimization
Deng Pan, Keerthiram Murugesan, Ting Hua +2
Understanding which parts of the retrieved context contribute to a large language model's generated answer is essential for building interpretable and trustworthy retrieval-augment…
cs.LG2026
Fast Explanations via Policy Gradient-Optimized Explainer
Deng Pan, Nuno Moniz, Nitesh Chawla
The challenge of delivering efficient explanations is a critical barrier that prevents the adoption of model explanations in real-world applications. Existing approaches often depe…