collaborators

5 papers

cs.CL2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin +315

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…

cs.CL2026

Agent-Driven Corpus Linguistics: A Framework for Autonomous Linguistic Discovery

Jia Yu, Weiwei Yu, Pengfei Xiao +1

Corpus linguistics has traditionally relied on human researchers to formulate hypotheses, construct queries, and interpret results - a process demanding specialized technical skill…

cs.CL2025

AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser

Ren Ma, Jiantao Qiu, Chao Xu +26

While web data quality is crucial for large language models, most curation efforts focus on filtering and deduplication,treating HTML-to-text extraction as a fixed pre-processing s…

cs.CL2025

Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering

Bowen Jiang, Runchuan Zhu, Jiang Wu +11

We introduce KoLasSimpleQA, the first benchmark evaluating the multilingual factual ability of Large Language Models (LLMs). Inspired by existing research, we created the question…

cs.CL2025

WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages

Jia Yu, Fei Yuan, Rui Min +20

This paper introduces the open-source dataset WanJuanSiLu, designed to provide high-quality training corpora for low-resource languages, thereby advancing the research and developm…