7 papers · 1 filter
An Iterative Algorithm for Rescaled Hyperbolic Functions Regression
Yeqi Gao, Zhao Song, Junze Yin
Large language models (LLMs) have numerous real-life applications across various domains, such as natural language translation, sentiment analysis, language modeling, chatbots and…
Efficient Alternating Minimization with Applications to Weighted Low Rank Approximation
Zhao Song, Mingquan Ye, Junze Yin +1
Weighted low rank approximation is a fundamental problem in numerical linear algebra, and it has many applications in machine learning. Given a matrix $M \in \mathbb{R}^{n \times n…
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
Yingyu Liang, Heshan Liu, Zhenmei Shi +3
The self-attention mechanism is the key to the success of transformers in recent Large Language Models (LLMs). However, the quadratic computational cost in the input seque…
Inverting the Leverage Score Gradient: An Efficient Approximate Newton Method
Chenyang Li, Zhao Song, Zhaoxing Xu +1
Leverage scores have become essential in statistics and machine learning, aiding regression analysis, randomized matrix computations, and various other tasks. This paper delves int…
How to Inverting the Leverage Score Distribution?
Zhihang Li, Zhao Song, Weixin Wang +2
Leverage score is a fundamental problem in machine learning and theoretical computer science. It has extensive applications in regression analysis, randomized algorithms, and neura…
Solving Attention Kernel Regression Problem via Pre-conditioner
Zhao Song, Junze Yin, Lichen Zhang
The attention mechanism is the key to large language models, and the attention matrix serves as an algorithmic and computational bottleneck for such a scheme. In this paper, we def…