6 papers
DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling
Yiming Ju, Hanyu Zhao, Quanyue Ma +5
Large-scale video repositories are increasingly available for modern video understanding and generation tasks. However, transforming raw videos into high-quality, task-specific dat…
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
Li Du, Hanyu Zhao, Yiming Ju +1
Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction…
AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning
Yifan Wei, Xiaoyan Yu, Yixuan Weng +3
Large Language Models (LLMs), when enhanced through reasoning-oriented post-training, evolve into powerful Large Reasoning Models (LRMs). Tool-Integrated Reasoning (TIR) further ex…
CI-VID: A Coherent Interleaved Text-Video Dataset
Yiming Ju, Jijin Hu, Zhengxiong Luo +7
Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this ar…
Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs
Yifan Wei, Xiaoyan Yu, Tengfei Pan +2
Large language models (LLMs) have achieved unprecedented performance by leveraging vast pretraining corpora, yet their performance remains suboptimal in knowledge-intensive domains…
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
Liangdong Wang, Bo-Wen Zhang, Chengwei Wu +7
We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/C…