5 papers
Dissociating the Internal Representations of Sycophancy in LLMs
Anthony Baez, Sheer Karny, Pat Pataranutaporn
Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophanc…
Guaranteeing Conservation of Integrals with Projection in Physics-Informed Neural Networks
Anthony Baez, Wang Zhang, Ziwen Ma +3
We propose a novel projection method that guarantees the conservation of integral quantities in Physics-Informed Neural Networks (PINNs). While the soft constraint that PINNs use t…
Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift
Sheer Karny, Anthony Baez, Pat Pataranutaporn
Chatbot behavior is often opaque to users, as responses can shift unpredictably across a conversation, drifting toward sycophancy, toxicity, or other unsafe responses. This can lea…
Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
Sheer Karny, Anthony Baez, Pat Pataranutaporn
Millions of users now design personalized LLM-based chatbots that shape their daily interactions, yet they can only roughly anticipate how their design choices will manifest as beh…
Guaranteeing Conservation Laws with Projection in Physics-Informed Neural Networks
Anthony Baez, Wang Zhang, Ziwen Ma +3
Physics-informed neural networks (PINNs) incorporate physical laws into their training to efficiently solve partial differential equations (PDEs) with minimal data. However, PINNs…