Publications (10)
Towards a theory of learning dynamics in deep state space models
Jakub Smékal, Jimmy T. H. Smith, Michael Kleinman +2
State space models (SSMs) have shown remarkable empirical performance on many long sequence modeling tasks, but a theoretical understanding of these models is still lacking. In thi…
LATTS: Locally Adaptive Test-Time Scaling
Theo Uscidda, Matthew Trager, Michael Kleinman +3
One common strategy for improving the performance of Large Language Models (LLMs) on downstream tasks involves using a \emph{verifier model} to either select the best answer from a…
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Renos Zabounidis, Aditya Golatkar, Michael Kleinman +3
We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking token…
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
Adam Stein, Matthew Trager, Benjamin Bowman +4
Enabling agentic AI systems to adapt their problem-solving approaches based on post-training interactions remains a fundamental challenge. While systems that update and maintain a…
e1: Learning Adaptive Control of Reasoning Effort
Michael Kleinman, Matthew Trager, Alessandro Achille +2
Increasing the thinking budget of AI models can significantly improve accuracy, but not all questions warrant the same amount of reasoning. Users may prefer to allocate different a…
Voice EHR: Introducing Multimodal Audio Data for Health
James Anibal, Hannah Huth, Ming Li +27
Artificial intelligence (AI) models trained on audio data may have the potential to rapidly perform clinical tasks, enhancing medical decision-making and potentially improving outc…
Critical Learning Periods Emerge Even in Deep Linear Networks
Michael Kleinman, Alessandro Achille, Stefano Soatto
Critical learning periods are periods early in development where temporary sensory deficits can have a permanent effect on behavior and learned representations. Despite the radical…
Usable Information and Evolution of Optimal Representations During Training
Michael Kleinman, Alessandro Achille, Daksh Idnani +1
We introduce a notion of usable information contained in the representation learned by a deep network, and use it to study how optimal representations for the task emerge during tr…
Gacs-Korner Common Information Variational Autoencoder
Michael Kleinman, Alessandro Achille, Stefano Soatto +1
We propose a notion of common information that allows one to quantify and separate the information that is shared between two random variables from the information that is unique t…
Critical Learning Periods for Multisensory Integration in Deep Networks
Michael Kleinman, Alessandro Achille, Stefano Soatto
We show that the ability of a neural network to integrate information from diverse sources hinges critically on being exposed to properly correlated signals during the early phases…