2 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.CL2024
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
Qingru Zhang, Xiaodong Yu, Chandan Singh +6
Large language models (LLMs) have demonstrated remarkable performance across various real-world tasks. However, they often struggle to fully comprehend and effectively utilize thei…
cs.CL2022★ 2 cited
MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation
Simiao Zuo, Qingru Zhang, Chen Liang +3
Pre-trained language models have demonstrated superior performance in various natural language processing tasks. However, these models usually contain hundreds of millions of param…
cs.LG2018
AdaShift: Decorrelation and Convergence of Adaptive Learning Rate Methods
Zhiming Zhou, Qingru Zhang, Guansong Lu +3
Adam is shown not being able to converge to the optimal solution in certain cases. Researchers recently propose several algorithms to avoid the issue of non-convergence of Adam, bu…