3 papers
cs.CL2026
Matryoshka attribution: Learning to attribute language model outputs to representations and weights
Aryaman Arora, Kirill Acharya, Nathan Hu +4
Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable…
cs.AI2026
Twin: Playing an Unknown Game with a Test-Time Digital Twin
Alexy Skoutnev, Kirill Acharya, Gaston Longhitano +3
We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-A…
math.OC2024
The Power of Extrapolation in Federated Learning
Hanmin Li, Kirill Acharya, Peter Richtárik
We propose and study several server-extrapolation strategies for enhancing the theoretical and empirical convergence properties of the popular federated learning optimizer FedProx…