most citedAttention Scheme Inspired Softmax Regression

6 citations · 18 across the 11 of their papers we have counts for

collaborators

11 papers

cs.LG2024

Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence

Yichuan Deng, Zhao Song, Chiwun Yang

Based on SGD, previous works have proposed many algorithms that have improved convergence speed and generalization in stochastic optimization, such as SGDm, AdaGrad, Adam, etc. How…

cs.LG20231 cited

Unmasking Transformers: A Theoretical Approach to Data Recovery via Attention Weights

Yichuan Deng, Zhao Song, Shenghao Xie +1

In the realm of deep learning, transformers have emerged as a dominant architecture, particularly in natural language processing tasks. However, with their widespread adoption, con…

cs.LG2023

Clustered Linear Contextual Bandits with Knapsacks

Yichuan Deng, Michalis Mamakos, Zhao Song

In this work, we study clustered contextual bandits where rewards and resource consumption are the outcomes of cluster-specific linear models. The arms are divided in clusters, wit…

cs.LG20232 cited

Convergence of Two-Layer Regression with Nonlinear Units

Yichuan Deng, Zhao Song, Shenghao Xie

Large language models (LLMs), such as ChatGPT and GPT4, have shown outstanding performance in many human life task. Attention computation plays an important role in training LLMs.…

cs.LG20233 cited

Zero-th Order Algorithm for Softmax Attention Optimization

Yichuan Deng, Zhihang Li, Sridhar Mahadevan +1

Large language models (LLMs) have brought about significant transformations in human society. Among the crucial computations in LLMs, the softmax unit holds great importance. Its h…

cs.LG20231 cited

Faster Robust Tensor Power Method for Arbitrary Order

Yichuan Deng, Zhao Song, Junze Yin

Tensor decomposition is a fundamental method used in various areas to deal with high-dimensional data. \emph{Tensor power method} (TPM) is one of the widely-used techniques in the…