3 papers
cs.AI2025
Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
Chengwei Wu, Li Du, Hanyu Zhao +4
Scaling the amount of data used for supervied fine-tuning(SFT) does not guarantee the proportional gains in model performance, highlighting a critical need to understand what makes…
cs.CV2025
CI-VID: A Coherent Interleaved Text-Video Dataset
Yiming Ju, Jijin Hu, Zhengxiong Luo +7
Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this ar…
cs.CL2024
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
Hanyu Zhao, Li Du, Yiming Ju +2
With the availability of various instruction datasets, a pivotal challenge is how to effectively select and integrate these instructions to fine-tune large language models (LLMs).…