13 citations · 16 across the 2 of their papers we have counts for
1 paper · 1 filter
Jiasen Lu, Christopher Clark, Sangho Lee +5
We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we…