Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers
Moritz Pawlowsky, Antonis Vamvakeros, Alexander Weiss +3
Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as…
cs.CV2025
Cube: A Roblox View of 3D Intelligence
Foundation AI Team, Kiran Bhat, Nishchaie Khanna +44
Foundation models trained on vast amounts of data have demonstrated remarkable reasoning and generation capabilities in the domains of text, images, audio and video. Our goal at Ro…