8 papers · 1 filter
When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili
Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains…
tensorFM: Low-Rank Approximations of Cross-Order Feature Interactions
Alessio Mazzetto, Mohammad Mahdi Khalili, Laura Fee Nern +3
We address prediction problems on tabular categorical data, where each instance is defined by multiple categorical attributes, each taking values from a finite set. These attribute…
Individual Fairness In Strategic Classification
Zhiqun Zuo, Mohammad Mahdi Khalili
Strategic classification, where individuals modify their features to influence machine learning (ML) decisions, presents critical fairness challenges. While group fairness in this…
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
Xudong Zhu, Mohammad Mahdi Khalili, Zhihui Zhu
Sparse autoencoders (SAEs) have emerged as powerful techniques for interpretability of large language models (LLMs), aiming to decompose hidden states into meaningful semantic feat…
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
Xudong Zhu, Jiachen Jiang, Mohammad Mahdi Khalili +1
Self-reflection -- the ability of a large language model (LLM) to revisit, evaluate, and revise its own reasoning -- has recently emerged as a powerful behavior enabled by reinforc…
Post-processing for Fair Regression via Explainable SVD
Zhiqun Zuo, Ding Zhu, Mohammad Mahdi Khalili
This paper presents a post-processing algorithm for training fair neural network regression models that satisfy statistical parity, utilizing an explainable singular value decompos…