1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2022
LaMemo: Language Modeling with Look-Ahead Memory
Haozhe Ji, Rongsheng Zhang, Zhenyu Yang +2
Although Transformers with fully connected self-attentions are powerful to model long-term dependencies, they are struggling to scale to long texts with thousands of words in langu…
cs.LG2021★ 1 cited
The distance between the weights of the neural network is meaningful
Liqun Yang, Yijun Yang, Yao Wang +2
In the application of neural networks, we need to select a suitable model based on the problem complexity and the dataset scale. To analyze the network's capacity, quantifying the…