3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2025
EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding
Mingxu Tao, Jie Hu, Mingchuan Yang +3
The remarkable performance of Large language models (LLMs) relies heavily on the availability of abundant high-quality training data. However, the high cost of acquiring annotated…
cs.LG2024★ 3 cited
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
Quzhe Huang, Zhenwei An, Nan Zhuang +7
In this paper, we introduce a novel dynamic expert selection framework for Mixture of Experts (MoE) models, aiming to enhance computational efficiency and model performance by adju…
cs.CL2023
Can BERT Refrain from Forgetting on Sequential Tasks? A Probing Study
Mingxu Tao, Yansong Feng, Dongyan Zhao
Large pre-trained language models help to achieve state of the art on a variety of natural language processing (NLP) tasks, nevertheless, they still suffer from forgetting when inc…