activity
20242026
most citedInternLM2 Technical Report

29 citations · 33 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser

Ren Ma, Jiantao Qiu, Chao Xu +26

While web data quality is crucial for large language models, most curation efforts focus on filtering and deduplication,treating HTML-to-text extraction as a fixed pre-processing s…

cs.CL2025

Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM

Mengjie Liu, Jiahui Peng, Wenchang Ning +14

High-quality main content extraction from web pages is a critical prerequisite for constructing large-scale training corpora. While traditional heuristic extractors are efficient,…

cs.CL2025

WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages

Jia Yu, Fei Yuan, Rui Min +20

This paper introduces the open-source dataset WanJuanSiLu, designed to provide high-quality training corpora for low-resource languages, thereby advancing the research and developm…

cs.CL2024

FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data

Haoran Sun, Renren Jin, Shaoyang Xu +10

Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource lan…

cs.CL202429 cited

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97

The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…

cs.CL2024

WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset

Jiantao Qiu, Haijun Lv, Zhenjiang Jin +23

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing larg…