6 papers
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
Carolina Zheng, Nicolas Beltran-Velez, Sweta Karlekar +5
Traditional topic models are effective at uncovering latent themes in large text collections. However, due to their reliance on bag-of-words representations, they struggle to captu…
Wrist Photoplethysmography Predicts Dietary Information
Kyle Verrier, Achille Nazaret, Joseph Futoma +2
Whether wearable photoplethysmography (PPG) contains dietary information remains unknown. We trained a language model on 1.1M meals to predict meal descriptions from PPG, aligning…
The CausalBench challenge: A machine learning contest for gene network inference from single-cell perturbation data
Mathieu Chevalley, Jacob Sackett-Sanders, Yusuf Roohani +15
In drug discovery, mapping interactions between genes within cellular systems is a crucial early step. Such maps are not only foundational for understanding the molecular mechanism…
Extremely Greedy Equivalence Search
Achille Nazaret, David Blei
The goal of causal discovery is to learn a directed acyclic graph from data. One of the most well-known methods for this problem is Greedy Equivalence Search (GES). GES searches fo…
Treeffuser: Probabilistic Predictions via Conditional Diffusions with Gradient-Boosted Trees
Nicolas Beltran-Velez, Alessandro Antonio Grande, Achille Nazaret +2
Probabilistic prediction aims to compute predictive distributions rather than single point predictions. These distributions enable practitioners to quantify uncertainty, compute ri…
Hypothesis Testing the Circuit Hypothesis in LLMs
Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret +5
Large language models (LLMs) demonstrate surprising capabilities, but we do not understand how they are implemented. One hypothesis suggests that these capabilities are primarily e…