activity
20242026
collaborators

7 papers

cs.LG2026

Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns

Vatsal Baherwani, Zixi Chen, Shikai Qiu +2

Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learn…

cs.LG2026

Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales

Shikai Qiu, Zixi Chen, Hoang Phan +2

Several recently introduced deep learning optimizers utilizing matrix-level preconditioning have shown promising speedups relative to the current dominant optimizer AdamW, particul…

cs.CL2026

OrthoGeoLoRA: Geometric Parameter-Efficient Fine-Tuning for Structured Social Science Concept Retrieval on theWeb

Zeqiang Wang, Xinyue Wu, Chenxi Li +4

Large language models and text encoders increasingly power web-based information systems in the social sciences, including digital libraries, data catalogues, and search interfaces…

math.OC2025

Convergence Rate in Nonlinear Two-Time-Scale Stochastic Approximation with State (Time)-Dependence

Zixi Chen, Yumin Xu, Ruixun Zhang

The nonlinear two-time-scale stochastic approximation is widely studied under conditions of bounded variances in noise. Motivated by recent advances that allow for variability link…

cond-mat.soft2025

A unifying approach to self-organizing systems interacting via conservation laws

Frank Barrows, Guanming Zhang, Satyam Anand +5

We present a unified framework for embedding and analyzing dynamical systems using generalized projection operators rooted in local conservation laws. By representing physical, bio…

cs.HC2025

Predicting Quality of Video Gaming Experience Using Global-Scale Telemetry Data and Federated Learning

Zhongyang Zhang, Jinhe Wen, Zixi Chen +6

Frames Per Second (FPS) significantly affects the gaming experience. Providing players with accurate FPS estimates prior to purchase benefits both players and game developers. Howe…