116 citations · 132 across the 7 of their papers we have counts for
1 paper · 1 filter
Euhid Aman, Esteban Carlin, Hsing-Kuo Pao +3
Cross-attention transformers and other multimodal vision-language models excel at grounding and generation; however, their extensive, full-precision backbones make it challenging t…