3 citations · 4 across the 10 of their papers we have counts for
11 papers · 1 filter
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
Dong Chen, Fangyun Wei, Ziyu Wan +18
We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters acro…
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
Yang Yue, Fangyun Wei, Tianyu He +10
Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on d…
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
Zeyu Liu, Zanlin Ni, Yang Yue +5
Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely deco…
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
Jiayi Guo, Linqing Wang, Jiangshan Wang +6
Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refi…
Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection
Jiaqi Wu, Zhen Wang, Enhao Huang +5
Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limit…
UltraSeP: Sequence-aware Pre-training for Echocardiography Probe Movement Guidance
Haojun Jiang, Teng Wang, Zhenguo Sun +8
Echocardiography is an essential medical technique for diagnosing cardiovascular diseases, but its high operational complexity has led to a shortage of trained professionals. To ad…