268 citations · 269 across the 3 of their papers we have counts for
3 papers
cs.CL2024★ 268 cited
DeepSeek-V3 Technical Report
DeepSeek-AI, Aixin Liu, Bei Feng +195
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effec…
cs.CL2024★ 1 cited
Flexible and Adaptable Summarization via Expertise Separation
Xiuying Chen, Mingzhe Li, Shen Gao +5
A proficient summarization model should exhibit both flexibility -- the capacity to handle a range of in-domain summarization tasks, and adaptability -- the competence to acquire n…
cs.LG2023
Regression with Cost-based Rejection
Xin Cheng, Yuzhou Cao, Haobo Wang +3
Learning with rejection is an important framework that can refrain from making predictions to avoid critical mispredictions by balancing between prediction and rejection. Previous…