collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2025

Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Jijie Li, Li Du, Hanyu Zhao +5

Large Language Models (LLMs) demonstrate strong performance in real-world applications, yet existing open-source instruction datasets often concentrate on narrow domains, such as m…

cs.CL2025

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models

Guang Liu, Liangdong Wang, Jijie Li +6

We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory. CCI4.0 occupies roughly TB…

cs.CL2025

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Shuhao Gu, Jialing Zhang, Siyuan Zhou +23

Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabiliti…

cs.CL2024

CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models

Liangdong Wang, Bo-Wen Zhang, Chengwei Wu +7

We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/C…

cs.CL2024

ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model

Shuhao Gu, Mengdi Zhao, Bowen Zhang +3

Tokenizer is an essential component for large language models (LLMs), and a tokenizer with a high compression rate can improve the model's representation and processing efficiency.…

cs.CL2024

Aquila2 Technical Report

Bo-Wen Zhang, Liangdong Wang, Jijie Li +6

This paper introduces the Aquila2 series, which comprises a wide range of bilingual models with parameter sizes of 7, 34, and 70 billion. These models are trained based on an innov…