1 paper · 1 filter
Sabbir Mollah, Rohit Gupta, Sirnam Swetha +3
Employing a single, unified model (UM) for both visual understanding (image-to-text: I2T) and visual generation (text-to-image: T2I) has opened a new direction in Visual Language M…