papers

Publications (15)

cs.CL2017

Video Highlight Prediction Using Audience Chat Reactions

Cheng-Yang Fu, Joon Lee, Mohit Bansal +1

Sports channel video portals offer an exciting domain for research on multimodal, multilingual analysis. We present methods addressing the problem of automatic video highlight pred…

cs.CV2023

FACET: Fairness in Computer Vision Evaluation Benchmark

Laura Gustafson, Chloe Rolland, Nikhila Ravi +5

Computer vision models have known performance disparities across attributes such as gender and skin tone. This means during tasks such as classification and detection, model perfor…

cs.CV2017

DSSD : Deconvolutional Single Shot Detector

Cheng-Yang Fu, Wei Liu, Ananth Ranga +2

The main contribution of this paper is an approach for introducing additional context into state-of-the-art general object detection. To achieve this we first combine a state-of-th…

cs.CV2015

Piecewise Linear Activation Functions For More Efficient Deep Networks

Cheng-Yang Fu, Alexander C. Berg

This submission has been withdrawn by arXiv administrators because it is intentionally incomplete, which is in violation of our policies.

cs.CV2019

RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free

Cheng-Yang Fu, Mykhailo Shvets, Alexander C. Berg

Recently two-stage detectors have surged ahead of single-shot detectors in the accuracy-vs-speed trade-off. Nevertheless single-shot detectors are immensely popular in embedded vis…

cs.CV2016

SSD: Single Shot MultiBox Detector

Wei Liu, Dragomir Anguelov, Dumitru Erhan +4

We present a method for detecting objects in images using a single deep neural network. Our approach, named SSD, discretizes the output space of bounding boxes into a set of defaul…

cs.CV2022

Hydra Attention: Efficient Attention with Many Heads

Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +2

While transformers have begun to dominate many tasks in vision, applying them to large images is still computationally difficult. A large reason for this is that self-attention sca…

cs.CV2022

Negative Frames Matter in Egocentric Visual Query 2D Localization

Mengmeng Xu, Cheng-Yang Fu, Yanghao Li +3

The recently released Ego4D dataset and benchmark significantly scales and diversifies the first-person visual perception data. In Ego4D, the Visual Queries 2D Localization task ai…

cs.CV2016

Fast Single Shot Detection and Pose Estimation

Patrick Poirson, Phil Ammirato, Cheng-Yang Fu +3

For applications in navigation and robotics, estimating the 3D pose of objects is as important as detection. Many approaches to pose estimation rely on detecting or tracking parts…

cs.CV2023

Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation

Tsu-Jui Fu, Licheng Yu, Ning Zhang +4

Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to…

cs.CV2019

IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things

Cheng-Yang Fu, Tamara L. Berg, Alexander C. Berg

In this work, we present a new operator, called Instance Mask Projection (IMP), which projects a predicted Instance Segmentation as a new feature for semantic segmentation. It also…

cs.CV2023

Token Merging: Your ViT But Faster

Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +3

We introduce Token Merging (ToMe), a simple method to increase the throughput of existing ViT models without needing to train. ToMe gradually combines similar tokens in a transform…

cs.CV2019

Target Driven Instance Detection

Phil Ammirato, Cheng-Yang Fu, Mykhailo Shvets +2

While state-of-the-art general object detectors are getting better and better, there are not many systems specifically designed to take advantage of the instance detection problem.…

cs.CV2022

End-to-End Visual Editing with a Generatively Pre-Trained Artist

Andrew Brown, Cheng-Yang Fu, Omkar Parkhi +2

We consider the targeted image editing problem: blending a region in a source image with a driver image that specifies the desired change. Differently from prior works, we solve th…

cs.CV2023

Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query Localization

Mengmeng Xu, Yanghao Li, Cheng-Yang Fu +3

This paper deals with the problem of localizing objects in image and video datasets from visual exemplars. In particular, we focus on the challenging problem of egocentric visual q…