activity
20232025
most citedInternLM2 Technical Report

29 citations · 42 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL2024

What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices

Zhi Chen, Qiguang Chen, Libo Qin +7

Recent advancements in large language models (LLMs) with extended context windows have significantly improved tasks such as information extraction, question answering, and complex…

cs.CL2024

Case2Code: Scalable Synthetic Data for Code Generation

Yunfan Shao, Linyang Li, Yichuan Ma +11

Large Language Models (LLMs) have shown outstanding breakthroughs in code generation. Recent work improves code LLMs by training on synthetic data generated by some powerful LLMs,…

cs.CL202429 cited

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97

The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…

cs.CL2024

WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset

Jiantao Qiu, Haijun Lv, Zhenjiang Jin +23

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing larg…

cs.CL2024

LongWanjuan: Towards Systematic Measurement for Long Text Quality

Kai Lv, Xiaoran Liu, Qipeng Guo +4

The quality of training data are crucial for enhancing the long-text capabilities of foundation models. Despite existing efforts to refine data quality through heuristic rules and…

cs.CL2024

Code Needs Comments: Enhancing Code LLMs with Comment Augmentation

Demin Song, Honglin Guo, Yunhua Zhou +8

The programming skill is one crucial ability for Large Language Models (LLMs), necessitating a deep understanding of programming languages (PLs) and their correlation with natural…