most citedHarnessing Diversity for Important Data Selection in Pretraining Large Language Models

3 citations · 3 across the 2 of their papers we have counts for

collaborators

5 papers

cs.DB2025

Unstructured Data Analysis using LLMs: A Comprehensive Benchmark

Qiyan Deng, Jianhui Li, Chengliang Chai +9

Nowadays, the explosion of unstructured data presents immense analytical value. Leveraging the remarkable capability of large language models (LLMs) in extracting attributes of str…

cs.DB2025

QUEST: Query Optimization in Unstructured Document Analysis

Zhaoze Sun, Qiyan Deng, Chengliang Chai +6

Most recently, researchers have started building large language models (LLMs) powered data systems that allow users to analyze unstructured text documents like working with a datab…

cs.CL2025

Not All Documents Are What You Need for Extracting Instruction Tuning Data

Chi Zhang, Huaping Zhong, Hongtao Li +11

Instruction tuning improves the performance of large language models (LLMs), but it heavily relies on high-quality training data. Recently, LLMs have been used to synthesize instru…

cs.LG2025

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

Kuan Zhang, Chengliang Chai, Jingzhe Xu +5

Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noi…

cs.AI20243 cited

Harnessing Diversity for Important Data Selection in Pretraining Large Language Models

Chi Zhang, Huaping Zhong, Kuan Zhang +10

Data selection is of great significance in pre-training large language models, given the variation in quality within the large-scale available training corpora. To achieve this, re…