4 papers
Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator
Li Hengyu
The attention matrix of a causal transformer is row-stochastic, iterated over depth, and non-normal by construction. For non-normal operators, eigenvalues control only asymptotic b…
Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention
Li Hengyu
The pre-softmax score of an attention head is a bilinear form in a learned operator . Because M is generally non-symmetric, hence non-norm…
Cross-layer Attention Sharing for Pre-trained Large Language Models
Yongyu Mu, Yuzhang Wu, Yuchun Fan +9
To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the KV cache or group attention heads, while largely ove…
Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
Yongyu Mu, Hengyu Li, Junxin Wang +7
Previous work on augmenting large multimodal models (LMMs) for text-to-image (T2I) generation has focused on enriching the input space of in-context learning (ICL). This includes p…