Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
High-Dimensional Robust Mean Estimation with Untrusted Batches
Maryam Aliakbarpour, Vladimir Braverman, Yuhan Liu +1
We study high-dimensional mean estimation in a collaborative setting where data is contributed by users in batches of size . In this environment, a learner seeks to recover…
cs.LG2025
Support Basis: Fast Attention Beyond Bounded Entries
Maryam Aliakbarpour, Vladimir Braverman, Junze Yin +1
Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks. However, the quadratic complexity of softmax attention remains a central bottlen…
cs.LG2025
Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
Haochen Zhang, Junze Yin, Guanchu Wang +5
Low-rank optimization has emerged as a promising approach to enabling memory-efficient training of large language models (LLMs). Existing low-rank optimization methods typically pr…