◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Pei Chu

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

fields
  • cs.CL3
  • cs.CV1

identity via Semantic Scholar / OpenAlex

most citedInternLM2 Technical Report

29 citations · 30 across the 4 of their papers we have counts for

collaborators

4 papers

cs.CL2025

WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages

Jia Yu, Fei Yuan, Rui Min +20

This paper introduces the open-source dataset WanJuanSiLu, designed to provide high-quality training corpora for low-resource languages, thereby advancing the research and developm…

cs.CV2024★ 1 cited

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Qingyun Li, Zhe Chen, Weiyun Wang +37

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resem…

cs.CL2024★ 29 cited

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97

The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…

cs.CL2024

WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset

Jiantao Qiu, Haijun Lv, Zhenjiang Jin +23

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing larg…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.