5 papers
LVFace: Progressive Cluster Optimization for Large Vision Models in Face Recognition
Jinghan You, Shanglin Li, Yuanrui Sun +4
Vision Transformers (ViTs) have revolutionized large-scale visual modeling, yet remain underexplored in face recognition (FR) where CNNs still dominate. We identify a critical bott…
ContentV: Efficient Training of Video Generation Models with Limited Compute
Wenfeng Lin, Renjie Chen, Boyuan Liu +10
Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter…
Towards Self-Improvement of Diffusion Models via Group Preference Optimization
Renjie Chen, Wenfeng Lin, Yichen Zhang +5
Aligning text-to-image (T2I) diffusion models with Direct Preference Optimization (DPO) has shown notable improvements in generation quality. However, applying DPO to T2I faces two…
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
Jiangchuan Wei, Shiyue Yan, Wenfeng Lin +3
Recent advancements in video generation have significantly impacted various downstream applications, particularly in identity-preserving video generation (IPT2V). However, existing…
CascadeV: An Implementation of Wurstchen Architecture for Video Generation
Wenfeng Lin, Jiangchuan Wei, Boyuan Liu +3
Recently, with the tremendous success of diffusion models in the field of text-to-image (T2I) generation, increasing attention has been directed toward their potential in text-to-v…