4 papers
On the Role of Preference Variance in Preference Optimization
Jiacheng Guo, Zihao Li, Jiahao Qiu +2
Direct Preference Optimization (DPO) has emerged as an important approach for learning from human preferences in aligning large language models (LLMs). However, collecting human pr…
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
Kaixuan Huang, Jiacheng Guo, Zihao Li +15
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieve…
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
Zihao Li, Yuan Cao, Cheng Gao +5
Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, the…
Global Convergence in Training Large-Scale Transformers
Cheng Gao, Yuan Cao, Zihao Li +5
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously an…