1 paper · 1 filter
Paul Pu Liang, Akshay Goindani, Talha Chafekar +4
Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applic…