4 papers
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI, Anyi Xu, Bangcai Lin +315
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…
CASTLE: Contrastive and Seed-Guided Training for Cold-Start Natural Language Search
Wendy Ran Wei, Hao Li, Weiwei Guo +9
Deploying natural language search systems presents a critical cold-start challenge: no real user queries to learn linguistic patterns, and no relevance labels to train ranking mode…
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
Haolei Bai, Lingcheng Kong, Xueyi Chen +3
Diffusion large language models (dLLMs) have emerged as a compelling alternative to autoregressive (AR) LLMs, owing to their capacity for parallel token generation. This paradigm i…
Training Hybrid Deep Quantum Neural Network for Efficient Reinforcement Learning
Jie Luo, Jeremy Kulcsar, Xueyin Chen +2
Quantum circuits embed data in a Hilbert space whose dimensionality grows exponentially with the number of qubits, allowing even shallow parameterised quantum circuits (PQCs) to re…