From the 1 of 7 linked papers with an AI index.
7 papers
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
Runhui Huang, Qihui Zhang, Zhe Liu +3
The paper introduces SpectraReward, a training-free method that uses pretrained multimodal large language models to score generated images by measuring how well the original text p…
Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
Team Seedance, Heyi Chen, Siyan Chen +194
Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically f…
Seedream 4.0: Toward Next-generation Multimodal Image Generation
Team Seedream, :, Yunpeng Chen +48
We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image compositi…
RewardDance: Reward Scaling in Visual Generation
Jie Wu, Yu Gao, Zilyu Ye +9
Reward Models (RMs) are critical for improving generation models via Reinforcement Learning (RL), yet the RM scaling paradigm in visual generation remains largely unexplored. It pr…
Seedance 1.0: Exploring the Boundaries of Video Generation Models
Yu Gao, Haoyuan Guo, Tuyen Hoang +41
Notable breakthroughs in diffusion modeling have propelled rapid improvements in video generation, yet current foundational model still face critical challenges in simultaneously b…
Seedream 3.0 Technical Report
Yu Gao, Lixue Gong, Qiushan Guo +28
We present Seedream 3.0, a high-performance Chinese-English bilingual image generation foundation model. We develop several technical improvements to address existing challenges in…