4 papers
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
Andrej Leban, Yuekai Sun
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely d…
Energy-Tweedie: Score meets Score, Energy meets Energy
Andrej Leban
Denoising and score estimation are classically linked through Tweedie's formula, which relates the posterior mean under Gaussian noise to the Stein score of the noisy marginal. In…
Distributional Autoencoders Know the Score
Andrej Leban
The Distributional Principal Autoencoder (DPA) combines distributionally correct reconstruction with principal-component-like interpretability of the encodings. In this work, we pr…
A Bayesian approach to translators' reliability assessment
Marco Miccheli, Andrej Leban, Andrea Tacchella +3
Translation Quality Assessment (TQA) is a process conducted by human translators and is widely used, both for estimating the performance of (increasingly used) Machine Translation,…