1 paper
Zhefan Rao, Bin Zou, Haoxuan Che +5
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-…