1 citations · 1 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
Yunlong Tang, Jing Bi, Chao Huang +16
We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects…
cs.CV2024
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
Zeliang Zhang, Zhuo Liu, Mingqian Feng +1
CLIP has demonstrated great versatility in adapting to various downstream tasks, such as image editing and generation, visual question answering, and video understanding. However,…
cs.CV2024★ 1 cited
Learning to Transform Dynamically for Better Adversarial Transferability
Rongyi Zhu, Zeliang Zhang, Susan Liang +2
Adversarial examples, crafted by adding perturbations imperceptible to humans, can deceive neural networks. Recent studies identify the adversarial transferability across various m…