3 papers
cs.LG2024
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
Hao Zhao, Zihan Qiu, Huijia Wu +3
The Mixture of Experts (MoE) for language models has been proven effective in augmenting the capacity of models by dynamically routing each input token to a specific subset of expe…
cs.AI2024
Read to Play (R2-Play): Decision Transformer with Multimodal Game Instruction
Yonggang Jin, Ge Zhang, Hao Zhao +7
Developing a generalist agent is a longstanding objective in artificial intelligence. Previous efforts utilizing extensive offline datasets from various tasks demonstrate remarkabl…
cs.CL2023
Prototype-based HyperAdapter for Sample-Efficient Multi-task Tuning
Hao Zhao, Jie Fu, Zhaofeng He
Parameter-efficient fine-tuning (PEFT) has shown its effectiveness in adapting the pre-trained language models to downstream tasks while only updating a small number of parameters.…