1 paper
Jaeik Kim, Woojin Kim, Jihwan Hong +8
We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understand…