3 papers
cs.CV2026
Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering
Yun Xing, Hanyuan Liu, Jiahao Nie +1
Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle region-level perception guide…
cs.CV2026
See-through: Single-image Layer Decomposition for Anime Characters
Jian Lin, Chengze Li, Haoyun Qin +5
We introduce a framework that automates the transformation of static anime illustrations into manipulatable 2.5D models. Current professional workflows require tedious manual segme…
cs.CV2025
Text-Guided Texturing by Synchronized Multi-View Diffusion
Yuxin Liu, Minshan Xie, Hanyuan Liu +1
This paper introduces a novel approach to synthesize texture to dress up a given 3D object, given a text prompt. Based on the pretrained text-to-image (T2I) diffusion model, existi…