1 citations · 1 across the 2 of their papers we have counts for
5 papers
Why Softmax Attention Outperforms Linear Attention
Yichuan Deng, Zhao Song, Kaijun Yuan +1
Large transformer models have achieved state-of-the-art results in numerous natural language processing tasks. Among the pivotal components of the transformer architecture, the att…
Dynamic Kernel Graph Sparsifiers
Yang Cao, Yichuan Deng, Wenyu Jin +4
A geometric graph associated with a set of points and a fixed kernel function $\mathsf{K}:\mathbb{R}^d\times \mathbb{R}^d\to\ma…
OpenThoughts: Data Recipes for Reasoning Models
Etash Guha, Ryan Marten, Sedrick Keh +47
Reasoning models have made rapid progress on many benchmarks involving math, code, and science. Yet, there are still many open questions about the best training recipes for reasoni…
Discrepancy Minimization in Input-Sparsity Time
Yichuan Deng, Xiaoyu Li, Zhao Song +1
A recent work by [Larsen, SODA 2023] introduced a faster combinatorial alternative to Bansal's SDP algorithm for finding a coloring that approximately minimizes…
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally -Sparse
Yichuan Deng, Zhao Song, Jing Xiong +1
Sparse Attention is a technique that approximates standard attention computation with sub-quadratic complexity. This is achieved by selectively ignoring smaller entries in the atte…