8 papers
MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
Yanhao Jia, Jiepeng Wang, Haibin Huang +3
Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation…
Style-CCL: Content-Preserving Style Transfer via Curriculum Continual Learning
Shiwen Zhang, Haoyuan Wang, Xianghao Zang +3
Content-Preserving Style transfer, given content and style references, remains challenging for Diffusion Transformers (DiTs) due to entangled content and style features. With a rev…
TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation
Shiwen Zhang, Yifan Xu, Haibin Huang +2
Given a content reference and a style reference, content-preserving style transfer requires the model to generate stylized outputs with content and style consistency. We introduced…
TeleStyle: Content-Preserving Style Transfer in Images and Videos
Shiwen Zhang, Xiaoyan Yang, Bojia Zi +3
Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the i…
The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
Dakuan Lu, Jiaqi Zhang, Cheng Yuan +2
Recent advances in large language models (LLMs) have been largely driven by scaling laws for individual models, which predict performance improvements as model parameters and data…
Theoretical Foundations of Scaling Law in Familial Models
Huan Song, Qingfei Zhao, Ting Long +4
Neural scaling laws have become foundational for optimizing large language model (LLM) training, yet they typically assume a single dense model output. This limitation effectively…