4 papers · 1 filter
InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion
Guangzhao Li, Qingyan Wei, Huayu Zheng +7
We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise learning from cross-category…
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Yi Xin, Qi Qin, Siqi Luo +29
We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by util…
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
Yan Tai, Luhao Zhu, Yunan Ding +4
Multimodal Large Language Models (MLLMs) demonstrate robust zero-shot capabilities across diverse vision-language tasks after training on mega-scale datasets. However, dense predic…
Link-Context Learning for Multimodal LLMs
Yan Tai, Weichen Fan, Zhao Zhang +3
The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…