1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2024
A Mean Field Ansatz for Zero-Shot Weight Transfer
Xingyuan Chen, Wenwei Kuang, Lei Deng +3
The pre-training cost of large language models (LLMs) is prohibitive. One cutting-edge approach to reduce the cost is zero-shot weight transfer, also known as model growth for some…
cs.CL2024★ 1 cited
Retrieval Meets Reasoning: Dynamic In-Context Editing for Long-Text Understanding
Weizhi Fei, Xueyan Niu, Guoqing Xie +4
Current Large Language Models (LLMs) face inherent limitations due to their pre-defined context lengths, which impede their capacity for multi-hop reasoning within extensive textua…
cs.LG2024
Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory
Xueyan Niu, Bo Bai, Lei Deng +1
Increasing the size of a Transformer does not always lead to enhanced performance. This phenomenon cannot be explained by the empirical scaling laws. Furthermore, the model's enhan…