8 papers
DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling
Yiming Ju, Hanyu Zhao, Quanyue Ma +5
Large-scale video repositories are increasingly available for modern video understanding and generation tasks. However, transforming raw videos into high-quality, task-specific dat…
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
Li Du, Hanyu Zhao, Yiming Ju +1
Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction…
Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
Chengwei Wu, Li Du, Hanyu Zhao +4
Scaling the amount of data used for supervied fine-tuning(SFT) does not guarantee the proportional gains in model performance, highlighting a critical need to understand what makes…
CI-VID: A Coherent Interleaved Text-Video Dataset
Yiming Ju, Jijin Hu, Zhengxiong Luo +7
Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this ar…
Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
Jijie Li, Li Du, Hanyu Zhao +5
Large Language Models (LLMs) demonstrate strong performance in real-world applications, yet existing open-source instruction datasets often concentrate on narrow domains, such as m…
Beyond the Lens: Quantifying the Impact of Scientific Documentaries through Amazon Reviews
Jill Naiman, Aria Pessianzadeh, Hanyu Zhao +8
Engaging the public with science is critical for a well-informed population. A popular method of scientific communication is documentaries. Once released, it can be difficult to as…