2 papers
cs.LG2026
Predicting LLM Safety Before Release by Simulating Deployment
Marcus Williams, Hannah Sheahan, Cameron Raymond +8
Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model beha…
cs.AI2025
Context Attribution with Multi-Armed Bandit Optimization
Deng Pan, Keerthiram Murugesan, Ting Hua +2
Understanding which parts of the retrieved context contribute to a large language model's generated answer is essential for building interpretable and trustworthy retrieval-augment…