3 papers
cs.CV2025
Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
Ciara Rowles, Varun Jampani, Simon Donné +3
Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We co…
cs.CV2024
Jointly Generating Multi-view Consistent PBR Textures using Collaborative Control
Shimon Vainer, Konstantin Kutsy, Dante De Nigris +3
Multi-view consistency remains a challenge for image diffusion models. Even within the Text-to-Texture problem, where perfect geometric correspondences are known a priori, many met…
cs.CV2024
IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning using Instruct Prompts
Ciara Rowles, Shimon Vainer, Dante De Nigris +3
Diffusion models continuously push the boundary of state-of-the-art image generation, but the process is hard to control with any nuance: practice proves that textual prompts are i…