4 papers
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
Cameron Braunstein, Mariya Toneva, Eddy Ilg
Latent diffusion models such as Stable Diffusion achieve state-of-the-art results on text-to-image generation tasks. However, the extent to which these models have a semantic under…
TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
Jonas Belouadi, Eddy Ilg, Margret Keuper +5
Automatically synthesizing figures from text captions is a compelling capability. However, achieving high geometric precision and editability requires representing figures as graph…
3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
Noor Ahmed, Cameron Braunstein, Steffen Eger +1
Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning rema…
Imaging for All-Day Wearable Smart Glasses
Michael Goesele, Daniel Andersen, Yujia Chen +7
In recent years smart glasses technology has rapidly advanced, opening up entirely new areas for mobile computing. We expect future smart glasses will need to be all-day wearable,…