activity
20222024
most citedMultimodal Speech Emotion Recognition using Cross Attention with Aligned Audio and Text

26 citations · 28 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck

Thang M. Pham, Peijie Chen, Tin Nguyen +3

CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rare…

cs.CV20241 cited

Scaling Up Video Summarization Pretraining with Large Language Models

Dawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron +5

Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summariza…

cs.CV2024

Fine-tuning CLIP Text Encoders with Two-step Paraphrasing

Hyunjae Kim, Seunghyun Yoon, Trung Bui +4

Contrastive language-image pre-training (CLIP) models have demonstrated considerable success across various vision-language tasks, such as text-to-image retrieval, where the model…

cs.CL20231 cited

Aspect-based Meeting Transcript Summarization: A Two-Stage Approach with Weak Supervision on Sentence Classification

Zhongfen Deng, Seunghyun Yoon, Trung Bui +7

Aspect-based meeting transcript summarization aims to produce multiple summaries, each focusing on one aspect of content in a meeting transcript. It is challenging as sentences rel…

cs.CL2023

Multilingual Sentence-Level Semantic Search using Meta-Distillation Learning

Meryem M'hamdi, Jonathan May, Franck Dernoncourt +2

Multilingual semantic search is the task of retrieving relevant contents to a query expressed in different language combinations. This requires a better semantic understanding of t…

cs.CL2023

Boosting Punctuation Restoration with Data Generation and Reinforcement Learning

Viet Dac Lai, Abel Salinas, Hao Tan +6

Punctuation restoration is an important task in automatic speech recognition (ASR) which aim to restore the syntactic structure of generated ASR texts to improve readability. While…