3 citations · 3 across the 17 of their papers we have counts for
1 paper · 1 filter
Puneet Mathur, Manan Suri, Dinesh Manocha
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent toke…