3 citations · 3 across the 11 of their papers we have counts for
8 papers · 1 filter
MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts
Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou +2
Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both re…
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
Sixun Dong, Juhua Hu, Mian Zhang +3
Vision-Language Models (VLMs) demonstrate impressive performance in understanding visual content with language instruction by converting visual inputs to vision tokens. However, re…
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
Qi Qian, Haiyang Xu, Ming Yan +1
Diffusion models demonstrate impressive image generation performance with text guidance. Inspired by the learning process of diffusion, existing images can be edited according to t…
Text-Guided Mixup Towards Long-Tailed Image Categorization
Richard Franklin, Jiawei Yao, Deyang Zhong +2
In many real-world applications, the frequency distribution of class labels for training data can exhibit a long-tailed distribution, which challenges traditional approaches of tra…
SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning
Qi Qian, Yuanhong Xu, Juhua Hu
Deep features extracted from certain layers of a pre-trained deep model show superior performance over the conventional hand-crafted features. Compared with fine-tuning or linear p…
Online Zero-Shot Classification with CLIP
Qi Qian, Juhua Hu
Vision-language pre-training such as CLIP enables zero-shot transfer that can classify images according to the candidate class names. While CLIP demonstrates an impressive zero-sho…