2 papers
cs.LG2026
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
Sungjun Lim, Heedong Kim, Andrew Lee +1
Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and…
cs.CL2025
Eeyore: Realistic Depression Simulation via Supervised and Preference Optimization
Siyang Liu, Bianca Brie, Wenda Li +4
Large Language Models (LLMs) have been previously explored for mental healthcare training and therapy client simulation, but they still fall short in authentically capturing divers…