1 citations · 1 across the 1 of their papers we have counts for
1 paper
Delong Chen, Mustafa Shukor, Theo Moutakanni +7
We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA…