collaborators

7 papers

cs.CL2025

Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models

Gabriele Prato, Shagun Sodhani, Alessandro Sordoni +1

The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on t…

cs.AI2025

AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench

Edan Toledo, Karen Hambardzumyan, Martin Josifoski +22

AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus o…

cs.CL2025

Do Large Language Models Know How Much They Know?

Gabriele Prato, Jerry Huang, Prasanna Parthasarathi +2

Large Language Models (LLMs) have emerged as highly capable systems and are increasingly being integrated into various uses. However, the rapid pace of their deployment has outpace…

eess.AS2025

Scaling and Distilling Transformer Models for sEMG

Nicholas Mehlman, Jean-Christophe Gagnon-Audet, Michael Shvartsman +3

Surface electromyography (sEMG) signals offer a promising avenue for developing innovative human-computer interfaces by providing insights into muscular activity. However, the limi…

cs.AI2025

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

Bingchen Zhao, Despoina Magka, Minqi Jiang +20

Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce exist…

cs.LG2025

Harnessing small projectors and multiple views for efficient vision pretraining

Kumar Krishna Agrawal, Arna Ghosh, Shagun Sodhani +2

Recent progress in self-supervised (SSL) visual representation learning has led to the development of several different proposed frameworks that rely on augmentations of images but…