6 papers
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
Peilin Xiong, Honghui Yuan, Junwen Chen +1
Coarse-mask local image editing asks a model to modify a user-indicated region while preserving the surrounding scene. In practice, however, rough masks often become unintended sha…
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
Qihua Dong, Ruozhen He, Junwen Chen +4
Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While existing MLLMs are strong at und…
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
Junwen Chen, Peilin Xiong, Keiji Yanai
Recent human-object interaction detection (HOID) methods highly require prior knowledge from vision-language models (VLMs) to enhance the interaction recognition capabilities. The…
PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
Peilin Xiong, Junwen Chen, Honghui Yuan +1
Localized subject-driven image editing aims to seamlessly integrate user-specified objects into target scenes. As generative models continue to scale, training becomes increasingly…
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
Junwen Chen, Heyang Jiang, Yanbin Wang +6
Generating high-quality, multi-layer transparent images from text prompts can unlock a new level of creative control, allowing users to edit each layer as effortlessly as editing t…
Focusing on what to decode and what to train: SOV Decoding with Specific Target Guided DeNoising and Vision Language Advisor
Junwen Chen, Yingcheng Wang, Keiji Yanai
Recent transformer-based methods achieve notable gains in the Human-object Interaction Detection (HOID) task by leveraging the detection of DETR and the prior knowledge of Vision-L…