papers

Publications (10)

cs.LG2024

Towards a theory of learning dynamics in deep state space models

Jakub Smékal, Jimmy T. H. Smith, Michael Kleinman +2

State space models (SSMs) have shown remarkable empirical performance on many long sequence modeling tasks, but a theoretical understanding of these models is still lacking. In thi…

cs.AI2025

LATTS: Locally Adaptive Test-Time Scaling

Theo Uscidda, Matthew Trager, Michael Kleinman +3

One common strategy for improving the performance of Large Language Models (LLMs) on downstream tasks involves using a \emph{verifier model} to either select the best answer from a…

cs.AI2025

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Renos Zabounidis, Aditya Golatkar, Michael Kleinman +3

We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking token…

cs.AI2025

Experience-Guided Adaptation of Inference-Time Reasoning Strategies

Adam Stein, Matthew Trager, Benjamin Bowman +4

Enabling agentic AI systems to adapt their problem-solving approaches based on post-training interactions remains a fundamental challenge. While systems that update and maintain a…

cs.AI2025

e1: Learning Adaptive Control of Reasoning Effort

Michael Kleinman, Matthew Trager, Alessandro Achille +2

Increasing the thinking budget of AI models can significantly improve accuracy, but not all questions warrant the same amount of reasoning. Users may prefer to allocate different a…

cs.SD2024

Voice EHR: Introducing Multimodal Audio Data for Health

James Anibal, Hannah Huth, Ming Li +27

Artificial intelligence (AI) models trained on audio data may have the potential to rapidly perform clinical tasks, enhancing medical decision-making and potentially improving outc…

cs.LG2024

Critical Learning Periods Emerge Even in Deep Linear Networks

Michael Kleinman, Alessandro Achille, Stefano Soatto

Critical learning periods are periods early in development where temporary sensory deficits can have a permanent effect on behavior and learned representations. Despite the radical…

cs.LG2021

Usable Information and Evolution of Optimal Representations During Training

Michael Kleinman, Alessandro Achille, Daksh Idnani +1

We introduce a notion of usable information contained in the representation learned by a deep network, and use it to study how optimal representations for the task emerge during tr…

cs.LG2023

Gacs-Korner Common Information Variational Autoencoder

Michael Kleinman, Alessandro Achille, Stefano Soatto +1

We propose a notion of common information that allows one to quantify and separate the information that is shared between two random variables from the information that is unique t…

cs.LG2023

Critical Learning Periods for Multisensory Integration in Deep Networks

Michael Kleinman, Alessandro Achille, Stefano Soatto

We show that the ability of a neural network to integrate information from diverse sources hinges critically on being exposed to properly correlated signals during the early phases…