works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.CL2026

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting

Mengya Hu, Susie Park, Suzana Ilic +5

The paper studies how to best place and combine content‑moderation filters and response rewriting in conversational systems, measuring overall usefulness and harmful exposure rathe…

cs.CV2026

Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds

Thomas Fel, Matthew Kowal, Mozes Jacobs +22

What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directio…

cs.LG2025

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Adam Karvonen, Can Rager, Johnny Lin +12

Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most pri…

cs.LG2025

Sparse Autoencoders Do Not Find Canonical Units of Analysis

Patrick Leask, Bart Bussmann, Michael Pearce +5

A common goal of mechanistic interpretability is to decompose the activations of neural networks into features: interpretable properties of the input computed by the model. Sparse…

cs.LG2024

LLM Circuit Analyses Are Consistent Across Training and Scale

Curt Tigges, Michael Hanna, Qinan Yu +1

Most currently deployed large language models (LLMs) undergo continuous training or additional finetuning. By contrast, most research into LLMs' internal mechanisms focuses on mode…