5 papers
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
Cameron Braunstein, Mariya Toneva, Eddy Ilg
Latent diffusion models such as Stable Diffusion achieve state-of-the-art results on text-to-image generation tasks. However, the extent to which these models have a semantic under…
The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends
Nico Policzer, Cameron Braunstein, Mariya Toneva
Recent studies on audio models show brain-tuning - fine-tuning models to better predict corresponding fMRI activity - improves brain alignment and increases performance on downstre…
3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
Noor Ahmed, Cameron Braunstein, Steffen Eger +1
Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning rema…
SLayR: Scene Layout Generation with Rectified Flow
Cameron Braunstein, Hevra Petekkaya, Jan Eric Lenssen +2
We introduce SLayR, Scene Layout Generation with Rectified flow, a novel transformer-based model for text-to-layout generation which can then be paired with existing layout-to-imag…
Quantum-Hybrid Stereo Matching With Nonlinear Regularization and Spatial Pyramids
Cameron Braunstein, Eddy Ilg, Vladislav Golyanik
Quantum visual computing is advancing rapidly. This paper presents a new formulation for stereo matching with nonlinear regularizers and spatial pyramids on quantum annealers as a…