Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Rongchang Xie, Chen Du, Ping Song +1
We introduce MUSE-VL, a Unified Vision-Language Model through Semantic discrete Encoding for multimodal understanding and generation. Recently, the research community has begun exp…
cs.CV2025
RefCut: Interactive Segmentation with Reference Guidance
Zheng Lin, Nan Zhou, Chen-Xi Du +2
Interactive segmentation aims to segment the specified target on the image with positive and negative clicks from users. Interactive ambiguity is a crucial issue in this field, whi…
cs.CV2025
Multi-Object Grounding via Hierarchical Contrastive Siamese Transformers
Chengyi Du, Keyan Jin
Multi-object grounding in 3D scenes involves localizing multiple objects based on natural language input. While previous work has primarily focused on single-object grounding, real…