activity
20222024
most citedLong-range Multimodal Pretraining for Movie Understanding

1 citations · 3 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Adapting Dual-encoder Vision-language Models for Paraphrased Retrieval

Jiacheng Cheng, Hijung Valentina Shin, Nuno Vasconcelos +2

In the recent years, the dual-encoder vision-language models (\eg CLIP) have achieved remarkable text-to-image retrieval performance. However, we discover that these models usually…

cs.CV2024

Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

Gihyun Kwon, Simon Jenni, Dingzeyu Li +3

While there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this…

cs.CV2024

Towards Automated Movie Trailer Generation

Dawit Mureja Argaw, Mattia Soldan, Alejandro Pardo +4

Movie trailers are an essential tool for promoting films and attracting audiences. However, the process of creating trailers can be time-consuming and expensive. To streamline this…

cs.CV20241 cited

Scaling Up Video Summarization Pretraining with Large Language Models

Dawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron +5

Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summariza…

cs.CV20231 cited

Long-range Multimodal Pretraining for Movie Understanding

Dawit Mureja Argaw, Joon-Young Lee, Markus Woodson +2

Learning computer vision models from (and for) movies has a long-standing history. While great progress has been attained, there is still a need for a pretrained multimodal model t…

cs.CV2022

The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video Editing

Dawit Mureja Argaw, Fabian Caba Heilbron, Joon-Young Lee +2

Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, colo…