6 papers
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries
Yuhui Wang, Jinqi Liu, Chengliang Chai +8
The diverse formats of CSV and Parquet files in data lakes pose a significant challenge to traditional ETL, which relies on data engineers to pre-define a target database schema an…
Unstructured Data Analysis using LLMs: A Comprehensive Benchmark
Qiyan Deng, Jianhui Li, Chengliang Chai +9
Nowadays, the explosion of unstructured data presents immense analytical value. Leveraging the remarkable capability of large language models (LLMs) in extracting attributes of str…
Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization
Kuan Zhang, Chengliang Chai, Jingzhe Xu +5
Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noi…
QUEST: Query Optimization in Unstructured Document Analysis
Zhaoze Sun, Qiyan Deng, Chengliang Chai +6
Most recently, researchers have started building large language models (LLMs) powered data systems that allow users to analyze unstructured text documents like working with a datab…
Not All Documents Are What You Need for Extracting Instruction Tuning Data
Chi Zhang, Huaping Zhong, Hongtao Li +11
Instruction tuning improves the performance of large language models (LLMs), but it heavily relies on high-quality training data. Recently, LLMs have been used to synthesize instru…
Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
Chi Zhang, Huaping Zhong, Kuan Zhang +10
Data selection is of great significance in pre-training large language models, given the variation in quality within the large-scale available training corpora. To achieve this, re…