activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

Sultan Alshehri, Zhantao Yang, Han Zhang +1

Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person"…

cs.CV2026

OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control

Yukun Wang, Ruihuang Li, Jiale Tao +7

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entan…

cs.CV2025

Addressing the ID-Matching Challenge in Long Video Captioning

Zhantao Yang, Huangji Wang, Ruili Feng +6

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal u…

cs.CV2025

MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance

Shangwen Zhu, Qianyu Peng, Zhilei Shu +9

High-fidelity text-to-image and text-to-video generation typically relies on Classifier-Free Guidance (CFG), but achieving optimal results often demands computationally expensive s…

cs.CV2025

Accelerating Diffusion Sampling via Exploiting Local Transition Coherence

Shangwen Zhu, Han Zhang, Zhantao Yang +4

Text-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the lengthy sampling time of the de…

cs.CV2024

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan +13

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…