5 papers · 1 filter
LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents
Haoyang Fang, Wei Zhu, Boran Han +11
RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization paramet…
Adapting to Online Distribution Shifts in Deep Learning: A Black-Box Approach
Dheeraj Baby, Boran Han, Shuai Zhang +3
We study the well-motivated problem of online distribution shift in which the data arrive in batches and the distribution of each batch can change arbitrarily over time. Since the…
Unraveling the Gradient Descent Dynamics of Transformers
Bingqing Song, Boran Han, Shuai Zhang +2
While the Transformer architecture has achieved remarkable success across various domains, a thorough theoretical foundation explaining its optimization dynamics is yet to be fully…
Transferring Knowledge from Large Foundation Models to Small Downstream Models
Shikai Qiu, Boran Han, Danielle C. Maddix +3
How do we transfer the relevant knowledge from ever larger foundation models into small, task-specific downstream models that can run at much lower costs? Standard transfer learnin…
Discovering Bias in Latent Space: An Unsupervised Debiasing Approach
Dyah Adila, Shuai Zhang, Boran Han +1
The question-answering (QA) capabilities of foundation models are highly sensitive to prompt variations, rendering their performance susceptible to superficial, non-meaning-alterin…