works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.LG2026

Bridging Compute- and Data-Optimal Pretraining

Tian Qin, Kimia Hamidieh, David Alvarez-Melis

The paper introduces Compute-Data (CD) scaling laws that unify compute-optimal and data-optimal pretraining regimes by modeling the effectiveness of derived tokens, and shows how t…

cs.LG2026

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

Rachit Bansal, Clara Mohri, Tian Qin +2

The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from…

cs.CV2026

Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution

Tian Qin, Junzhe Chen, Yuqing Shi +3

Large vision-language models (LVLMs) often hallucinate when language priors dominate weak or ambiguous visual evidence. Existing contrastive decoding methods mitigate this problem…

cs.LG2026

Drawback of Enforcing Equivariance and its Compensation via the Lens of Expressive Power

Yuzhu Chen, Tian Qin, Xinmei Tian +2

Equivariant neural networks encode the intrinsic symmetry of data as an inductive bias, which has achieved impressive performance in wide domains. However, the understanding to the…

cs.LG2026

Random Scaling of Emergent Capabilities

Rosie Zhao, Tian Qin, David Alvarez-Melis +2

Language models famously improve under a smooth scaling law, but some specific capabilities exhibit sudden breakthroughs in performance. Advocates of "emergence" view these capabil…

cs.LG2025

To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning

Tian Qin, David Alvarez-Melis, Samy Jelassi +1

Recent advancements in large language models (LLMs) have significantly improved their reasoning abilities, particularly through techniques involving search and backtracking. Backtr…