works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.LG2026

Bridging Compute- and Data-Optimal Pretraining

Tian Qin, Kimia Hamidieh, David Alvarez-Melis

The paper introduces Compute-Data (CD) scaling laws that unify compute-optimal and data-optimal pretraining regimes by modeling the effectiveness of derived tokens, and shows how t…

cs.AI2026

Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3

Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design problems to be difficult for humans or c…

cs.LG2026

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

Rachit Bansal, Clara Mohri, Tian Qin +2

The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from…

cs.CV2026

Low-Frequency Shortcuts in Texture-Driven Visual Learning

Utku Şirin, Cathy Hou, David Alvarez-Melis +1

Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets. Ex…

cs.LG2026

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Jing Huang, Daniel Wurgaft, Rachit Bansal +6

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger mo…

cs.CY2026

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

Saleh Afroogh, Syed Ishtiaque Ahmed, Petra Ahrweiler +46

This study provides a cross-disciplinary examination of Explainable Artificial Intelligence (XAI) approaches-focusing on deep neural networks (DNNs) and large language models (LLMs…