2 citations · 2 across the 13 of their papers we have counts for
21 papers · 1 filter
LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter
Tobias Christian Nauen, Anosh Billimoria, Federico Raue +3
Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the b…
OA-CutMix: Correcting the Label Bias of CutMix
Tobias Christian Nauen, Stanislav Frolov, Federico Raue +2
CutMix has become the de facto standard mixing augmentation, yet its label assignment rests on a flawed assumption: The area of the pasted patch faithfully reflects its semantic co…
TextTeacher: What Can Language Teach About Images?
Tobias Christian Nauen, Stanislav Frolov, Brian Bernhard Moser +3
The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Motivated by this, we ask: Can…
LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency
Ryugo Morita, Stanislav Frolov, Brian Bernhard Moser +4
Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances in layered image generation a…
LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation
Ryugo Morita, Stanislav Frolov, Brian Bernhard Moser +3
Diffusion models have demonstrated high-quality performance in conditional text-to-image generation, particularly with structural cues such as edges, layouts, and depth. However, l…
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
Krzysztof Adamkiewicz, Brian Bernhard Moser, Stanislav Frolov +3
Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as synthetic vision data generator…