works on

From the 1 of 12 linked papers with an AI index.

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Bridging Compute- and Data-Optimal Pretraining

Tian Qin, Kimia Hamidieh, David Alvarez-Melis

The paper introduces Compute-Data (CD) scaling laws that unify compute-optimal and data-optimal pretraining regimes by modeling the effectiveness of derived tokens, and shows how t…

cs.LG2026

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

Rachit Bansal, Clara Mohri, Tian Qin +2

The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from…

cs.LG2026

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Jing Huang, Daniel Wurgaft, Rachit Bansal +6

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger mo…

cs.LG2026

Random Scaling of Emergent Capabilities

Rosie Zhao, Tian Qin, David Alvarez-Melis +2

Language models famously improve under a smooth scaling law, but some specific capabilities exhibit sudden breakthroughs in performance. Advocates of "emergence" view these capabil…

cs.LG2026

Reliable and Responsible Foundation Models: A Comprehensive Survey

Xinyu Yang, Junlin Han, Rishi Bommasani +49

Foundation models, including Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), Image Generative Models (i.e, Text-to-Image Models and Image-Editing Models), a…

cs.LG2025

To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning

Tian Qin, David Alvarez-Melis, Samy Jelassi +1

Recent advancements in large language models (LLMs) have significantly improved their reasoning abilities, particularly through techniques involving search and backtracking. Backtr…