6 citations · 14 across the 10 of their papers we have counts for
11 papers
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
Weizhi Wang, Khalil Mrini, Linjie Yang +4
We propose a novel framework for filtering image-text data by leveraging fine-tuned Multimodal Language Models (MLMs). Our approach outperforms predominant filtering methods (e.g.,…
Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling
Haogeng Liu, Qihang Fan, Tingkai Liu +5
This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text…
Selective Feature Adapter for Dense Vision Transformers
Xueqing Deng, Qi Fan, Xiaojie Jin +2
Fine-tuning pre-trained transformer models, e.g., Swin Transformer, are successful in numerous downstream for dense prediction vision tasks. However, one major issue is the cost/st…
The Devil is in the Details: A Deep Dive into the Rabbit Hole of Data Filtering
Haichao Yu, Yu Tian, Sateesh Kumar +2
The quality of pre-training data plays a critical role in the performance of foundation models. Popular foundation models often design their own recipe for data filtering, which ma…
Learning Dynamic Query Combinations for Transformer-based Object Detection and Segmentation
Yiming Cui, Linjie Yang, Haichao Yu
Transformer-based detection and segmentation methods use a list of learned detection queries to retrieve information from the transformer network and learn to predict the location…
Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?
Cheng-En Wu, Yu Tian, Haichao Yu +4
Vision-language models such as CLIP learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through…