6 papers
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
Peilin Xiong, Honghui Yuan, Junwen Chen +1
Coarse-mask local image editing asks a model to modify a user-indicated region while preserving the surrounding scene. In practice, however, rough masks often become unintended sha…
Point-MF: One-step Point Cloud Generation from a Single Image via Mean Flows
Yuta Baba, Keiji Yanai
Single-image point cloud reconstruction must infer complete 3D geometry, including occluded parts, from a single RGB image. While diffusion-based reconstructors achieve high accura…
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
Keisuke Gomi, Keiji Yanai
Cross-modal retrieval between food images and recipe texts is an important task with applications in nutritional management, dietary logging, and cooking assistance. Existing metho…
SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
Honghui Yuan, Keiji Yanai
With the rapid development of diffusion models, style transfer has made remarkable progress. However, flexible and localized style editing for scene text remains an unsolved challe…
PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
Peilin Xiong, Junwen Chen, Honghui Yuan +1
Localized subject-driven image editing aims to seamlessly integrate user-specified objects into target scenes. As generative models continue to scale, training becomes increasingly…
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
Junwen Chen, Heyang Jiang, Yanbin Wang +6
Generating high-quality, multi-layer transparent images from text prompts can unlock a new level of creative control, allowing users to edit each layer as effortlessly as editing t…