Showing 2026Show all
3 papers · 1 filter
cs.CV2026
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
Peilin Xiong, Honghui Yuan, Junwen Chen +1
Coarse-mask local image editing asks a model to modify a user-indicated region while preserving the surrounding scene. In practice, however, rough masks often become unintended sha…
cs.CV2026
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
Qihua Dong, Ruozhen He, Junwen Chen +4
Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While existing MLLMs are strong at und…
cs.CV2026
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
Junwen Chen, Peilin Xiong, Keiji Yanai
Recent human-object interaction detection (HOID) methods highly require prior knowledge from vision-language models (VLMs) to enhance the interaction recognition capabilities. The…