2 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Shaolei Zhang, Shoutao Guo, Qingkai Fang +2
The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal intera…