1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 1 cited
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong +7
The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and…
cs.CL2023
End-to-end Story Plot Generator
Hanlin Zhu, Andrew Cohen, Danqing Wang +4
Story plots, while short, carry most of the essential information of a full story that may contain tens of thousands of words. We study the problem of automatic generation of story…
cs.LG2023
Sample-efficient Surrogate Model for Frequency Response of Linear PDEs using Self-Attentive Complex Polynomials
Andrew Cohen, Weiping Dou, Jiang Zhu +7
Linear Partial Differential Equations (PDEs) govern the spatial-temporal dynamics of physical systems that are essential to building modern technology. When working with linear PDE…