collaborators

8 papers

cs.CL2026

Automated Attribution Graph Interpretation via Probe Prompting

Giuseppe Birardi, Gonçalo Paulo

Even though we know the precise computations that lead from a large language model (LLM) input to its output this computation remains very hard to interpret. One way to make it eas…

physics.chem-ph2025

Electrodrying in nanopores: from fundamentals to iontronic and memristive applications

Giovanni Di Muccio, Gonçalo Paulo, Lorenzo Iannetti +3

Iontronics is a burgeoning paradigm that employs ions in solution as information carriers for sensing and computing, e.g., in neuromorphic devices. The fundamentally different work…

cs.LG2025

Automatically Interpreting Millions of Features in Large Language Models

Gonçalo Paulo, Alex Mallen, Caden Juang +1

While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these…

cs.LG2025

Evaluating SAE interpretability without explanations

Gonçalo Paulo, Nora Belrose

Sparse autoencoders (SAEs) and transcoders have become important tools for machine learning interpretability. However, measuring how interpretable they are remains challenging, wit…

cs.CL2025

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Guijin Son, Jiwoo Hong, Honglu Fan +8

Recent advances in large language models (LLMs) have fueled the vision of automated scientific discovery, often called AI Co-Scientists. To date, prior work casts these systems as…

cs.LG2025

Transcoders Beat Sparse Autoencoders for Interpretability

Gonçalo Paulo, Stepan Shabalin, Nora Belrose

Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…