7 papers
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
Gabriele Prato, Shagun Sodhani, Alessandro Sordoni +1
The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on t…
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Edan Toledo, Karen Hambardzumyan, Martin Josifoski +22
AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus o…
Do Large Language Models Know How Much They Know?
Gabriele Prato, Jerry Huang, Prasanna Parthasarathi +2
Large Language Models (LLMs) have emerged as highly capable systems and are increasingly being integrated into various uses. However, the rapid pace of their deployment has outpace…
Scaling and Distilling Transformer Models for sEMG
Nicholas Mehlman, Jean-Christophe Gagnon-Audet, Michael Shvartsman +3
Surface electromyography (sEMG) signals offer a promising avenue for developing innovative human-computer interfaces by providing insights into muscular activity. However, the limi…
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Bingchen Zhao, Despoina Magka, Minqi Jiang +20
Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce exist…
Harnessing small projectors and multiple views for efficient vision pretraining
Kumar Krishna Agrawal, Arna Ghosh, Shagun Sodhani +2
Recent progress in self-supervised (SSL) visual representation learning has led to the development of several different proposed frameworks that rely on augmentations of images but…