1 paper
Tong Zhao, Yuyang Hu, Yutao Zhu +5
Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data…