75 citations · 99 across the 5 of their papers we have counts for
5 papers
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI, Anyi Xu, Bangcai Lin +315
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…
A Universal Discriminator for Zero-Shot Generalization
Haike Xu, Zongyu Lin, Jing Zhou +2
Generative modeling has been the dominant approach for large-scale pretraining and zero-shot generalization. In this work, we challenge this convention by showing that discriminati…
FewNLU: Benchmarking State-of-the-Art Methods for Few-Shot Natural Language Understanding
Yanan Zheng, Jing Zhou, Yujie Qian +7
The few-shot natural language understanding (NLU) task has attracted much recent attention. However, prior methods have been evaluated under a disparate set of protocols, which hin…
FlipDA: Effective and Robust Data Augmentation for Few-Shot Learning
Jing Zhou, Yanan Zheng, Jie Tang +2
Most previous methods for text data augmentation are limited to simple tasks and weak baselines. We explore data augmentation on hard tasks (i.e., few-shot natural language underst…
SYSTRAN's Pure Neural Machine Translation Systems
Josep Crego, Jungi Kim, Guillaume Klein +27
Since the first online demonstration of Neural Machine Translation (NMT) by LISA, NMT development has recently moved from laboratory to production systems as demonstrated by severa…