activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

metabeta -- A fast neural model for Bayesian mixed-effects regression

Alex Kipnis, Marcel Binz, Eric Schulz

Hierarchical data with multiple observations per group is ubiquitous in empirical sciences and is often analyzed using mixed-effects regression. In such models, Bayesian inference…

cs.LG2025

Exploring System 1 and 2 communication for latent reasoning in LLMs

Julian Coda-Forno, Zhuokai Zhao, Qiang Zhang +6

Should LLM reasoning live in a separate module, or within a single model's forward pass and representational space? We study dual-architecture latent reasoning, where a fluent Base…

cs.LG2025

A circuit for predicting hierarchical structure in-context in Large Language Models

Tankred Saanum, Can Demircan, Samuel J. Gershman +1

Large Language Models (LLMs) excel at in-context learning, the ability to use information provided as context to improve prediction of future tokens. Induction heads have been argu…

cs.LG2025

Automated scientific minimization of regret

Marcel Binz, Akshay K. Jagadish, Milena Rmus +1

We introduce automated scientific minimization of regret (ASMR) -- a framework for automated computational cognitive science. Building on the principles of scientific regret minimi…

cs.LG2024

Simplifying Latent Dynamics with Softly State-Invariant World Models

Tankred Saanum, Peter Dayan, Eric Schulz

To solve control problems via model-based reasoning or planning, an agent needs to know how its actions affect the state of the world. The actions an agent has at its disposal ofte…

cs.LG2024

Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models

Can Demircan, Tankred Saanum, Akshay K. Jagadish +2

In-context learning, the ability to adapt based on a few examples in the input prompt, is a ubiquitous feature of large language models (LLMs). However, as LLMs' in-context learnin…