2 papers
cs.CL2024
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies
Bo-Wen Zhang, Liangdong Wang, Ye Yuan +24
In recent years, with the rapid application of large language models across various fields, the scale of these models has gradually increased, and the resources required for their…
cs.CL2023
Aligner: One Global Token is Worth Millions of Parameters When Aligning Large Language Models
Zhou Ziheng, Yingnian Wu, Song-Chun Zhu +1
We introduce Aligner, a novel Parameter-Efficient Fine-Tuning (PEFT) method for aligning multi-billion-parameter-sized Large Language Models (LLMs). Aligner employs a unique design…