activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

UniFusion: Vision-Language Model as Unified Encoder in Image Generation

Kevin Li, Manuel Brack, Sudeep Katakol +2

Although recent advances in visual generation have been remarkable, most existing architectures still depend on distinct encoders for images and text. This separation constrains di…

cs.CV2025

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions

Manuel Brack, Sudeep Katakol, Felix Friedrich +4

Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inc…

cs.CV2025

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Lukas Helff, Felix Friedrich, Manuel Brack +2

This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models. To this e…

cs.CV2024

Core Tokensets for Data-efficient Sequential Training of Transformers

Subarnaduti Paul, Manuel Brack, Patrick Schramowski +2

Deep networks are frequently tuned to novel tasks and continue learning from ongoing data streams. Such sequential training requires consolidation of new and past information, a ch…

cs.CV2024

LEDITS++: Limitless Image Editing using Text-to-Image Models

Manuel Brack, Felix Friedrich, Katharina Kornmeier +4

Text-to-image diffusion models have recently received increasing interest for their astonishing ability to produce high-fidelity images from solely text inputs. Subsequent research…