3 papers
cs.CV2026
What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers
Moritz Pawlowsky, Antonis Vamvakeros, Alexander Weiss +3
Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as…
cs.CV2025
Cube: A Roblox View of 3D Intelligence
Foundation AI Team, Kiran Bhat, Nishchaie Khanna +44
Foundation models trained on vast amounts of data have demonstrated remarkable reasoning and generation capabilities in the domains of text, images, audio and video. Our goal at Ro…
cs.GR2024
FlashTex: Fast Relightable Mesh Texturing with LightControlNet
Kangle Deng, Timothy Omernick, Alexander Weiss +4
Manually creating textures for 3D meshes is time-consuming, even for expert visual content creators. We propose a fast approach for automatically texturing an input 3D mesh based o…