1 citations · 2 across the 3 of their papers we have counts for
3 papers
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
Chaofan Tao, Gukyeong Kwon, Varad Gunjal +7
We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes pa…
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
Luca Zancato, Arjun Seshadri, Yonatan Dukler +6
We describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resou…
Learning Expressive Prompting With Residuals for Vision Transformers
Rajshekhar Das, Yonatan Dukler, Avinash Ravichandran +1
Prompt learning is an efficient approach to adapt transformers by inserting learnable set of parameters into the input and intermediate representations of a pre-trained model. In t…