1 paper · 1 filter
Jaeik Kim, Woojin Kim, Jihwan Hong +8
We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understand…