3 citations · 3 across the 2 of their papers we have counts for
6 papers
Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
Qiyan Deng, Changqian Zheng, Lianpeng Qiao +3
Dataset distillation condenses large datasets into synthetic subsets, achieving performance comparable to training on the full dataset while substantially reducing storage and comp…
Unstructured Data Analysis using LLMs: A Comprehensive Benchmark
Qiyan Deng, Jianhui Li, Chengliang Chai +9
Nowadays, the explosion of unstructured data presents immense analytical value. Leveraging the remarkable capability of large language models (LLMs) in extracting attributes of str…
QUEST: Query Optimization in Unstructured Document Analysis
Zhaoze Sun, Qiyan Deng, Chengliang Chai +6
Most recently, researchers have started building large language models (LLMs) powered data systems that allow users to analyze unstructured text documents like working with a datab…
Not All Documents Are What You Need for Extracting Instruction Tuning Data
Chi Zhang, Huaping Zhong, Hongtao Li +11
Instruction tuning improves the performance of large language models (LLMs), but it heavily relies on high-quality training data. Recently, LLMs have been used to synthesize instru…
Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization
Kuan Zhang, Chengliang Chai, Jingzhe Xu +5
Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noi…
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
Chi Zhang, Huaping Zhong, Kuan Zhang +10
Data selection is of great significance in pre-training large language models, given the variation in quality within the large-scale available training corpora. To achieve this, re…