2 citations · 2 across the 9 of their papers we have counts for
10 papers
When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon
Huaqing Zhang, Jingchu Gai, Juno Kim +2
Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT). Y…
Momentum Streams for Optimizer-Inspired Transformers
Jingchu Gai, Nai-Chieh Huang, Jiayun Wu
The residual update of a pre-norm Transformer layer admits an interpretation as one step of a first-order optimizer acting on a surrogate token energy, wherein the attention and ML…
Lossless Anti-Distillation Sampling
Zibo Diao, Jingchu Gai, Xinyue Ai +3
Frontier commercial generative models face a growing threat from distillation, whereby a distiller harvests generated responses and trains a competing model of its own at drastical…
Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation
Jingchu Gai, Laixi Shi
Multi-agent reinforcement learning (MARL) holds great potential but faces robustness challenges due to environmental uncertainty. To address this, distributionally robust Markov ga…
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
Jingchu Gai, Guanning Zeng, Christina Baek +4
Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving rea…
Towards Solving the Gilbert-Pollak Conjecture via Large Language Models
Yisi Ke, Tianyu Huang, Yankai Shu +3
The Gilbert-Pollak Conjecture \citep{gilbert1968steiner}, also known as the Steiner Ratio Conjecture, states that for any finite point set in the Euclidean plane, the Steiner minim…