7 papers · 1 filter
Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models
Sultan Alshehri, Zhantao Yang, Han Zhang +1
Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person"…
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
Yukun Wang, Ruihuang Li, Jiale Tao +7
Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entan…
Addressing the ID-Matching Challenge in Long Video Captioning
Zhantao Yang, Huangji Wang, Ruili Feng +6
Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal u…
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
Shangwen Zhu, Qianyu Peng, Zhilei Shu +9
High-fidelity text-to-image and text-to-video generation typically relies on Classifier-Free Guidance (CFG), but achieving optimal results often demands computationally expensive s…
Accelerating Diffusion Sampling via Exploiting Local Transition Coherence
Shangwen Zhu, Han Zhang, Zhantao Yang +4
Text-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the lengthy sampling time of the de…
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
Zhantao Yang, Ruili Feng, Keyu Yan +13
Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…