4 papers
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
Aditya Nawal, Manit Baser, Mohan Gurusamy
AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the gene…
CLaRE-ty Amid Chaos: Quantifying Representational Entanglement to Predict Ripple Effects in LLM Editing
Manit Baser, Alperen Yildiz, Dinil Mon Divakaran +1
The static knowledge representations of large language models (LLMs) inevitably become outdated or incorrect over time. While model-editing techniques offer a promising solution by…
ThinkEval: Practical Evaluation of Knowledge Leakage in LLM Editing using Thought-based Knowledge Graphs
Manit Baser, Dinil Mon Divakaran, Mohan Gurusamy
Robust model-editing techniques are essential for deploying large language models (LLMs) in practical applications, as they enable cost-effective ways to deal with challenges such…
Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
Yash Sinha, Manit Baser, Murari Mandal +2
Knowledge erasure in large language models (LLMs) is important for ensuring compliance with data and AI regulations, safeguarding user privacy, mitigating bias, and misinformation.…