2 papers
cs.LG2026
QK-Normed MLA: QK normalization without full key caching
Yizhou Han, Yao Zhao, Jun Zhou +2
Query-key (QK) normalization stabilizes attention by controlling the scale of queries and keys before the dot product, but is not immediately compatible with Multi-head Latent Atte…
cs.LG2026
DriftGuard: Mitigating Asynchronous Data Drift in Federated Learning
Yizhou Han, Di Wu, Blesson Varghese
In real-world Federated Learning (FL) deployments, data distributions on devices that participate in training evolve over time. This leads to asynchronous data drift, where differe…