activity
20182025
most citedVIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

90 citations · 248 across the 21 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

cs.CV2021★ 90 cited

VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Tsu-Jui Fu, Linjie Li, Zhe Gan +4

A great challenge in video-language (VidL) modeling lies in the disconnection between fixed video representations extracted from image/video understanding models and downstream Vid…

cs.CV2021

Language-Driven Image Style Transfer

Tsu-Jui Fu, Xin Eric Wang, William Yang Wang

Despite having promising results, style transfer, which requires preparing style images in advance, may result in lack of creativity and accessibility. Following human instruction,…

cs.CV2021

M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers

Tsu-Jui Fu, Xin Eric Wang, Scott T. Grafton +2

Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get sta…

cs.CV2021★ 2 cited

L2C: Describing Visual Differences Needs Semantic Understanding of Individuals

An Yan, Xin Eric Wang, Tsu-Jui Fu +1

Recent advances in language and vision push forward the research of captioning a single image to describing visual differences between image pairs. Suppose there are two images, I_…

cs.CV2021

DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents

Tsu-Jui Fu, William Yang Wang, Daniel McDuff +1

Creating presentation materials requires complex multimodal reasoning skills to summarize key concepts and arrange them in a logical and visually pleasing manner. Can machines lear…