3 citations · 3 across the 4 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning
Haoqiang Kang, Yinpeng Chen, Luyang Liu +5
Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a…
cs.CV2023
Disentangling the Effects of Data Augmentation and Format Transform in Self-Supervised Learning of Image Representations
Neha Kalibhat, Warren Morningstar, Alex Bijamov +3
Self-Supervised Learning (SSL) enables training performant models using limited labeled data. One of the pillars underlying vision SSL is the use of data augmentations/perturbation…