2 citations · 2 across the 24 of their papers we have counts for
18 papers · 1 filter
Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment
Zihao Wang, Xi Xiang, Yuwen Sun +5
Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, tex…
Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer
Bo Wei, Xianhui Lin, Yi Dong +8
Makeup transfer applies a reference cosmetic style to a source face while preserving its identity and geometry. However, this task is severely hindered by the lack of real paired t…
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
Renlong Wu, Haoran Chen, Yuxiang Wei +3
Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D…
InstanceControl: Controllable Complex Image Generation without Instance Labeling
Xiaoyu Liu, Huan Wang, Fan Li +4
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. Howev…
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Chonghuinan Wang, Zhikai Chen, Chunwei Wang +9
The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving…
Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning
Mengzhao Wang, Yanli Ji, Wangmeng Zuo +2
Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeated…