activity
20242026
collaborators

6 papers

cs.DB2026

EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries

Yuhui Wang, Jinqi Liu, Chengliang Chai +8

The diverse formats of CSV and Parquet files in data lakes pose a significant challenge to traditional ETL, which relies on data engineers to pre-define a target database schema an…

cs.DB2025

Unstructured Data Analysis using LLMs: A Comprehensive Benchmark

Qiyan Deng, Jianhui Li, Chengliang Chai +9

Nowadays, the explosion of unstructured data presents immense analytical value. Leveraging the remarkable capability of large language models (LLMs) in extracting attributes of str…

cs.LG2025

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

Kuan Zhang, Chengliang Chai, Jingzhe Xu +5

Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noi…

cs.DB2025

QUEST: Query Optimization in Unstructured Document Analysis

Zhaoze Sun, Qiyan Deng, Chengliang Chai +6

Most recently, researchers have started building large language models (LLMs) powered data systems that allow users to analyze unstructured text documents like working with a datab…

cs.CL2025

Not All Documents Are What You Need for Extracting Instruction Tuning Data

Chi Zhang, Huaping Zhong, Hongtao Li +11

Instruction tuning improves the performance of large language models (LLMs), but it heavily relies on high-quality training data. Recently, LLMs have been used to synthesize instru…

cs.AI2024

Harnessing Diversity for Important Data Selection in Pretraining Large Language Models

Chi Zhang, Huaping Zhong, Kuan Zhang +10

Data selection is of great significance in pre-training large language models, given the variation in quality within the large-scale available training corpora. To achieve this, re…