1 paper
Chao Feng, Zihao Wei, Andrew Owens
We learn visual features by captioning images with an image-conditioned masked diffusion language model, a formulation we call masked diffusion captioning (MDC). During training, t…