1 paper
Fan Bai, Pai Peng, Zhengzhi Tang +8
With the widespread adoption of large multimodal models, efficient inference across text, image, audio, and video modalities has become critical. However, existing multimodal infer…