3 papers
cs.CL2025
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
Xin Gao, Xiaoyang Wang, Yun Zhu +3
The construction of Supervised Fine-Tuning (SFT) datasets is a critical yet under-theorized stage in the post-training of Large Language Models (LLMs), as prevalent practices often…
cs.AI2025
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
Mengzhang Cai, Xin Gao, Yu Li +13
The rapid evolution of Large Language Models (LLMs) is predicated on the quality and diversity of post-training datasets. However, a critical dichotomy persists: while models are r…
cs.AI2025
LLM/Agent-as-Data-Analyst: A Survey
Zirui Tang, Weizheng Wang, Zihang Zhou +16
Large language models (LLMs) and agent techniques have brought a fundamental shift in the functionality and development paradigm of data analysis tasks (a.k.a LLM/Agent-as-Data-Ana…