1 paper
Han Yang, Kun Su, Yutong Zhang +4
We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To ad…