26 citations · 28 across the 8 of their papers we have counts for
8 papers
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
Thang M. Pham, Peijie Chen, Tin Nguyen +3
CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rare…
Scaling Up Video Summarization Pretraining with Large Language Models
Dawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron +5
Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summariza…
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
Hyunjae Kim, Seunghyun Yoon, Trung Bui +4
Contrastive language-image pre-training (CLIP) models have demonstrated considerable success across various vision-language tasks, such as text-to-image retrieval, where the model…
Aspect-based Meeting Transcript Summarization: A Two-Stage Approach with Weak Supervision on Sentence Classification
Zhongfen Deng, Seunghyun Yoon, Trung Bui +7
Aspect-based meeting transcript summarization aims to produce multiple summaries, each focusing on one aspect of content in a meeting transcript. It is challenging as sentences rel…
Multilingual Sentence-Level Semantic Search using Meta-Distillation Learning
Meryem M'hamdi, Jonathan May, Franck Dernoncourt +2
Multilingual semantic search is the task of retrieving relevant contents to a query expressed in different language combinations. This requires a better semantic understanding of t…
Boosting Punctuation Restoration with Data Generation and Reinforcement Learning
Viet Dac Lai, Abel Salinas, Hao Tan +6
Punctuation restoration is an important task in automatic speech recognition (ASR) which aim to restore the syntactic structure of generated ASR texts to improve readability. While…