268 citations · 269 across the 4 of their papers we have counts for
4 papers
DeepSeek-V3 Technical Report
DeepSeek-AI, Aixin Liu, Bei Feng +195
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effec…
Towards a Unified View of Preference Learning for Large Language Models: A Survey
Bofei Gao, Feifan Song, Yibo Miao +22
Large Language Models (LLMs) exhibit remarkably powerful capabilities. One of the crucial factors to achieve success is aligning the LLM's output with human preferences. This align…
Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
Zihan Wang, Peiyi Wang, Houfeng Wang
Hierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex taxonomic structure. Nearly all recent HTC works focus on how the l…
RepCL: Exploring Effective Representation for Continual Text Classification
Yifan Song, Peiyi Wang, Dawei Zhu +3
Continual learning (CL) aims to constantly learn new knowledge over time while avoiding catastrophic forgetting on old tasks. In this work, we focus on continual text classificatio…