1 paper
Xin He, Longhui Wei, Jianbo Ouyang +3
We propose EMMA, an efficient and unified architecture for multimodal understanding, generation and editing. Specifically, EMMA primarily consists of 1) An efficient autoencoder wi…