10 papers
A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula
Chenruo Liu, Yijun Dong, Yiqiu Shen +1
Iterative self-improvement fine-tunes an autoregressive large language model (LLM) on reward-verified outputs generated by the LLM itself. In contrast to the empirical success of s…
Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension
Yijun Dong, Yicheng Li, Yunai Li +2
Weak-to-strong (W2S) generalization is a type of finetuning (FT) where a strong (large) student model is trained on pseudo-labels generated by a weak teacher. Surprisingly, W2S FT…
Attention Mechanisms Through the Lens of Numerical Methods: Approximation Methods and Alternative Formulations
Michel Fabrice Serret, Alice Cortinovis, Yijun Dong +10
The attention mechanism is the computational core of modern Transformer architectures, but its quadratic complexity in the input sequence length is the bottleneck for large-scale i…
Does Weak-to-strong Generalization Happen under Spurious Correlations?
Chenruo Liu, Yijun Dong, Qi Lei
We initiate a unified theoretical and algorithmic study of a key problem in weak-to-strong (W2S) generalization: when fine-tuning a strong pre-trained student with pseudolabels fro…
When does Chain-of-Thought Help: A Markovian Perspective
Zihan Wang, Yijun Dong, Qi Lei
Chain-of-Thought (CoT) prompting is a widely used inference-time technique for improving reasoning, yet its gains are uneven across tasks. We analyze when and why CoT helps by mode…
Randomized time stepping of nonlinearly parametrized solutions of evolution problems
Yijun Dong, Paul Schwerdtner, Benjamin Peherstorfer
The Dirac-Frenkel variational principle is a widely used building block for using nonlinear parametrizations in the context of model reduction and numerically solving partial diffe…