1 paper
Tong Zhao, Yuyang Hu, Reed Li +5
Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data…