5 papers
DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling
Yiming Ju, Hanyu Zhao, Quanyue Ma +5
Large-scale video repositories are increasingly available for modern video understanding and generation tasks. However, transforming raw videos into high-quality, task-specific dat…
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
Li Du, Hanyu Zhao, Yiming Ju +1
Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction…
Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
Chengwei Wu, Li Du, Hanyu Zhao +4
Scaling the amount of data used for supervied fine-tuning(SFT) does not guarantee the proportional gains in model performance, highlighting a critical need to understand what makes…
CI-VID: A Coherent Interleaved Text-Video Dataset
Yiming Ju, Jijin Hu, Zhengxiong Luo +7
Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this ar…
Training Data for Large Language Model
Yiming Ju, Huanhuan Ma
In 2022, with the release of ChatGPT, large-scale language models gained widespread attention. ChatGPT not only surpassed previous models in terms of parameters and the scale of it…