activity
20182026
most citedERNIE: Enhanced Language Representation with Informative Entities

135 citations · 281 across the 29 of their papers we have counts for

collaborators

33 papers

cs.CL2026

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, :, Anyi Xu +585

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation…

cs.CL2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin +315

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…

cs.CL2024

Exploring the Benefit of Activation Sparsity in Pre-training

Zhengyan Zhang, Chaojun Xiao, Qiujieli Qin +7

Pre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activatio…

cs.AI2024★ 4 cited

Configurable Foundation Models: Building LLMs from a Modular Perspective

Chaojun Xiao, Zhengyan Zhang, Chenyang Song +20

Advancements in LLMs have recently unveiled challenges tied to computational efficiency and continual scalability due to their requirements of huge parameters, making the applicati…

cs.CL2024★ 1 cited

Robust and Scalable Model Editing for Large Language Models

Yingfa Chen, Zhengyan Zhang, Xu Han +6

Large language models (LLMs) can make predictions using parametric knowledge--knowledge encoded in the model weights--or contextual knowledge--knowledge presented in the context. I…

cs.LG2024★ 3 cited

ReLU Wins: Discovering Efficient Activation Functions for Sparse LLMs

Zhengyan Zhang, Yixin Song, Guanghui Yu +7

Sparse computation offers a compelling solution for the inference of Large Language Models (LLMs) in low-resource scenarios by dynamically skipping the computation of inactive neur…