5 papers
RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
Tianyuan Qu, Lei Ke, Xiaohang Zhan +6
The paper presents RePlan, a framework that first reasons about natural‑language instructions to identify specific image regions and then edits those regions using a diffusion mode…
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
Hanxun Yu, Xuan Qu, Lei Ke +4
Despite advances in 3D scene understanding, existing 3D Large Multimodal Models operate in offline settings, requiring complete scene observations or predefined video clips. In thi…
AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning
Peifeng Zhang, Zice Qiu, Donghua Yu +4
In continual visual question answering (VQA), existing Continual Learning (CL) methods are mostly built for symmetric, unimodal architectures. However, modern Vision-Language Model…
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
Boqiang Zhang, Lei Ke, Ruihan Yang +5
Vision Language Model (VLM) development has largely relied on scaling model size, which hinders deployment on compute-constrained mobile and edge devices such as smartphones and ro…
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
Yuxin Wang, Lei Ke, Boqiang Zhang +6
While current multimodal models can answer questions based on 2D images, they lack intrinsic 3D object perception, limiting their ability to comprehend spatial relationships and de…