collaborators

5 papers

cs.LG2026

HyperTransfer: Understanding the Equivalence between Base Optimizer and Hyperball

Jinghui Yuan, Hongtao Zhang, Jade Zou +4

Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. Although this geometry appears fu…

cs.LG2026

The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

Hongtao Zhang, Wenjie Zhou, Chenxi Jia +2

Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We identify an underlying spectral p…

cs.LG2026

Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training

Wenjie Zhou, Bohan Wang, Hongtao Zhang +3

Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze lat…

cs.IR2026

Benchmarking Real-Time Question Answering via Executable Code Workflows

Wenjie Zhou, Yuan Gao, Xin Zhou +5

Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and ther…

cs.LG2026

When and Why Grouping Attention Heads Accelerates Muon Optimization

Hongtao Zhang, Wenjie Zhou, Wei Chen +1

Muon orthogonalizes matrix updates, but multi-head attention naturally operates at the level of heads. This granularity mismatch raises the question of whether Muon should be appli…