activity
20242026
collaborators

6 papers

cs.CL2026

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders

Carolina Zheng, Nicolas Beltran-Velez, Sweta Karlekar +5

Traditional topic models are effective at uncovering latent themes in large text collections. However, due to their reliance on bag-of-words representations, they struggle to captu…

cs.LG2025

Wrist Photoplethysmography Predicts Dietary Information

Kyle Verrier, Achille Nazaret, Joseph Futoma +2

Whether wearable photoplethysmography (PPG) contains dietary information remains unknown. We trained a language model on 1.1M meals to predict meal descriptions from PPG, aligning…

cs.LG2025

The CausalBench challenge: A machine learning contest for gene network inference from single-cell perturbation data

Mathieu Chevalley, Jacob Sackett-Sanders, Yusuf Roohani +15

In drug discovery, mapping interactions between genes within cellular systems is a crucial early step. Such maps are not only foundational for understanding the molecular mechanism…

cs.LG2025

Extremely Greedy Equivalence Search

Achille Nazaret, David Blei

The goal of causal discovery is to learn a directed acyclic graph from data. One of the most well-known methods for this problem is Greedy Equivalence Search (GES). GES searches fo…

cs.LG2024

Treeffuser: Probabilistic Predictions via Conditional Diffusions with Gradient-Boosted Trees

Nicolas Beltran-Velez, Alessandro Antonio Grande, Achille Nazaret +2

Probabilistic prediction aims to compute predictive distributions rather than single point predictions. These distributions enable practitioners to quantify uncertainty, compute ri…

cs.AI2024

Hypothesis Testing the Circuit Hypothesis in LLMs

Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret +5

Large language models (LLMs) demonstrate surprising capabilities, but we do not understand how they are implemented. One hypothesis suggests that these capabilities are primarily e…