8 papers
Automated Attribution Graph Interpretation via Probe Prompting
Giuseppe Birardi, Gonçalo Paulo
Even though we know the precise computations that lead from a large language model (LLM) input to its output this computation remains very hard to interpret. One way to make it eas…
Electrodrying in nanopores: from fundamentals to iontronic and memristive applications
Giovanni Di Muccio, Gonçalo Paulo, Lorenzo Iannetti +3
Iontronics is a burgeoning paradigm that employs ions in solution as information carriers for sensing and computing, e.g., in neuromorphic devices. The fundamentally different work…
Automatically Interpreting Millions of Features in Large Language Models
Gonçalo Paulo, Alex Mallen, Caden Juang +1
While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these…
Evaluating SAE interpretability without explanations
Gonçalo Paulo, Nora Belrose
Sparse autoencoders (SAEs) and transcoders have become important tools for machine learning interpretability. However, measuring how interpretable they are remains challenging, wit…
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Guijin Son, Jiwoo Hong, Honglu Fan +8
Recent advances in large language models (LLMs) have fueled the vision of automated scientific discovery, often called AI Co-Scientists. To date, prior work casts these systems as…
Transcoders Beat Sparse Autoencoders for Interpretability
Gonçalo Paulo, Stepan Shabalin, Nora Belrose
Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…