11 papers
Constraint-based Pre-training: From Structured Constraints to Scalable Model Initialization
Fu Feng, Yucheng Xie, Ruixiao Shi +2
The pre-training and fine-tuning paradigm has become the dominant approach for model adaptation. However, conventional pre-training typically yields models at a fixed scale, wherea…
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
Yusheng Zheng, Wenan Mao, Shuyi Cheng +8
Performance diagnosis in production-scale AI training is challenging because subtle OS-level issues can trigger cascading GPU delays and network slowdowns, degrading training effic…
A Creative Agent is Worth a 64-Token Template
Ruixiao Shi, Fu Feng, Yucheng Xie +3
Text-to-image (T2I) models have substantially improved image fidelity and prompt adherence, yet their creativity remains constrained by reliance on discrete natural language prompt…
A Unified Framework for Knowledge Transfer in Bidirectional Model Scaling
Jianlu Shen, Fu Feng, Jiaze Xu +3
Transferring pre-trained knowledge from a source model to a target model of a different architectural size is a key challenge for flexible and efficient model scaling. However, cur…
Self-Supervised Weight Templates for Scalable Vision Model Initialization
Yucheng Xie, Fu Feng, Ruixiao Shi +3
The increasing scale and complexity of modern model parameters underscore the importance of pre-trained models. However, deployment often demands architectures of varying sizes, ex…
Knowledge Diversion for Efficient Morphology Control and Policy Transfer
Fu Feng, Ruixiao Shi, Yucheng Xie +3
Universal morphology control aims to learn a universal policy that generalizes across heterogeneous agent morphologies, with Transformer-based controllers emerging as a popular cho…