57 citations · 168 across the 19 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
Keer Lu, Xiaonan Nie, Zheng Liang +8
In recent years, Large Language Models (LLMs) have demonstrated significant improvements across a variety of tasks, one of which is the long-context capability. The key to improvin…
cs.CL2024
Data Proportion Detection for Optimized Data Management for Large Language Models
Hao Liang, Keshi Zhao, Yajie Yang +4
Large language models (LLMs) have demonstrated exceptional performance across a wide range of tasks and domains, with data preparation playing a critical role in achieving these re…