activity
20242026
collaborators

6 papers

cs.CV2026

DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling

Yiming Ju, Hanyu Zhao, Quanyue Ma +5

Large-scale video repositories are increasingly available for modern video understanding and generation tasks. However, transforming raw videos into high-quality, task-specific dat…

cs.AI2026

Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report

Li Du, Hanyu Zhao, Yiming Ju +1

Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction…

cs.CL2025

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning

Yifan Wei, Xiaoyan Yu, Yixuan Weng +3

Large Language Models (LLMs), when enhanced through reasoning-oriented post-training, evolve into powerful Large Reasoning Models (LRMs). Tool-Integrated Reasoning (TIR) further ex…

cs.CV2025

CI-VID: A Coherent Interleaved Text-Video Dataset

Yiming Ju, Jijin Hu, Zhengxiong Luo +7

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this ar…

cs.CL2025

Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs

Yifan Wei, Xiaoyan Yu, Tengfei Pan +2

Large language models (LLMs) have achieved unprecedented performance by leveraging vast pretraining corpora, yet their performance remains suboptimal in knowledge-intensive domains…

cs.CL2024

CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models

Liangdong Wang, Bo-Wen Zhang, Chengwei Wu +7

We present CCI3.0-HQ (https://huggingface.co/datasets/BAAI/CCI3-HQ), a high-quality 500GB subset of the Chinese Corpora Internet 3.0 (CCI3.0)(https://huggingface.co/datasets/BAAI/C…