1 paper
Chunlei Meng, Pengbin Feng, Rong Fu +7
Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers…