From the 1 of 12 linked papers with an AI index.
12 papers
Bridging Compute- and Data-Optimal Pretraining
Tian Qin, Kimia Hamidieh, David Alvarez-Melis
The paper introduces Compute-Data (CD) scaling laws that unify compute-optimal and data-optimal pretraining regimes by modeling the effectiveness of derived tokens, and shows how t…
Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry
Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3
Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design problems to be difficult for humans or c…
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training
Rachit Bansal, Clara Mohri, Tian Qin +2
The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from…
Low-Frequency Shortcuts in Texture-Driven Visual Learning
Utku Åirin, Cathy Hou, David Alvarez-Melis +1
Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets. Ex…
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
Jing Huang, Daniel Wurgaft, Rachit Bansal +6
Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger mo…
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
Saleh Afroogh, Syed Ishtiaque Ahmed, Petra Ahrweiler +46
This study provides a cross-disciplinary examination of Explainable Artificial Intelligence (XAI) approaches-focusing on deep neural networks (DNNs) and large language models (LLMs…