collaborators

6 papers

cs.AI2026

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Boyan Li, Zhuowen Liang, Yupeng Xie +11

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multi…

cs.LG2026

TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins

Yuxiang Luo, Haonan Long, Chen Wang +6

Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and naïve runs can…

cs.CL2026

Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs

Zhuowen Liang, Xiaotian Lin, Zhengxuan Zhang +3

Large language models (LLMs) are widely applied to data analytics over documents, yet direct reasoning over long, noisy documents remains brittle and error-prone. Hence, we study d…

cs.DB2026

A Survey of Data Agents: Emerging Paradigm or Overstated Hype?

Yizhang Zhu, Liangwei Wang, Chenyu Yang +22

The rapid advancement of large language models (LLMs) has spurred the emergence of data agents, autonomous systems designed to orchestrate Data + AI ecosystems for tackling complex…

cs.AI2025

Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting

Yifan Wu, Jingze Shi, Bingheng Wu +4

Existing chain-of-thought (CoT) distillation methods can effectively transfer reasoning abilities to base models but suffer from two major limitations: excessive verbosity of reaso…

cs.LG2025

LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning

Xiaotian Lin, Yanlin Qi, Yizhang Zhu +4

Instruction tuning has emerged as a critical paradigm for improving the capabilities and alignment of large language models (LLMs). However, existing iterative model-aware data sel…