4 papers · 1 filter
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
Ming Tao, Bing-Kun Bao, Yaowei Wang +1
Large pretrained diffusion models have demonstrated impressive generation capabilities and have been adapted to various downstream tasks. However, unlike Large Language Models (LLM…
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
Linhui Xiao, Xiaoshan Yang, Fang Peng +2
Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a v…
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
Linhui Xiao, Xiaoshan Yang, Fang Peng +2
Visual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained m…
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
Ming Tao, Bing-Kun Bao, Hao Tang +2
Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trai…