2 papers
cs.CV2026
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
Jiayun Luo, Mir Rayat Imtiaz Hossain, Pritam Sarkar +2
Vision-Language Models (VLMs) have achieved strong performance on implicit and explicit visual grounding and related tasks. However, such abilities are generally tested on simple,…
cs.CV2026
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
Xiangteng He, Shunsuke Sakai, Shivam Chandhok +5
Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In p…