2 papers
cs.LG2026
A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
Claire O'Brien, Jessica Seto, Dristi Roy +6
Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low inter…
cs.LG2025
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
Sri Vatsa Vuddanti, Aarav Shah, Satwik Kumar Chittiprolu +4
Tool-augmented language agents frequently fail in real-world deployment due to tool malfunctions--timeouts, API exceptions, or inconsistent outputs--triggering cascading reasoning…