Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Advancing Expert Specialization for Better MoE
Hongcan Guo, Haolang Lu, Guoshun Nan +8
Mixture-of-Experts (MoE) models enable efficient scaling of large language models (LLMs) by activating only a subset of experts per input. However, we observe that the commonly use…
cs.CL2023
DocMSU: A Comprehensive Benchmark for Document-level Multimodal Sarcasm Understanding
Hang Du, Guoshun Nan, Sicheng Zhang +6
Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks an…