Publications (15)
Video Highlight Prediction Using Audience Chat Reactions
Cheng-Yang Fu, Joon Lee, Mohit Bansal +1
Sports channel video portals offer an exciting domain for research on multimodal, multilingual analysis. We present methods addressing the problem of automatic video highlight pred…
FACET: Fairness in Computer Vision Evaluation Benchmark
Laura Gustafson, Chloe Rolland, Nikhila Ravi +5
Computer vision models have known performance disparities across attributes such as gender and skin tone. This means during tasks such as classification and detection, model perfor…
DSSD : Deconvolutional Single Shot Detector
Cheng-Yang Fu, Wei Liu, Ananth Ranga +2
The main contribution of this paper is an approach for introducing additional context into state-of-the-art general object detection. To achieve this we first combine a state-of-th…
Piecewise Linear Activation Functions For More Efficient Deep Networks
Cheng-Yang Fu, Alexander C. Berg
This submission has been withdrawn by arXiv administrators because it is intentionally incomplete, which is in violation of our policies.
RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
Cheng-Yang Fu, Mykhailo Shvets, Alexander C. Berg
Recently two-stage detectors have surged ahead of single-shot detectors in the accuracy-vs-speed trade-off. Nevertheless single-shot detectors are immensely popular in embedded vis…
SSD: Single Shot MultiBox Detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan +4
We present a method for detecting objects in images using a single deep neural network. Our approach, named SSD, discretizes the output space of bounding boxes into a set of defaul…
Hydra Attention: Efficient Attention with Many Heads
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +2
While transformers have begun to dominate many tasks in vision, applying them to large images is still computationally difficult. A large reason for this is that self-attention sca…
Negative Frames Matter in Egocentric Visual Query 2D Localization
Mengmeng Xu, Cheng-Yang Fu, Yanghao Li +3
The recently released Ego4D dataset and benchmark significantly scales and diversifies the first-person visual perception data. In Ego4D, the Visual Queries 2D Localization task ai…
Fast Single Shot Detection and Pose Estimation
Patrick Poirson, Phil Ammirato, Cheng-Yang Fu +3
For applications in navigation and robotics, estimating the 3D pose of objects is as important as detection. Many approaches to pose estimation rely on detecting or tracking parts…
Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation
Tsu-Jui Fu, Licheng Yu, Ning Zhang +4
Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to…
IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
Cheng-Yang Fu, Tamara L. Berg, Alexander C. Berg
In this work, we present a new operator, called Instance Mask Projection (IMP), which projects a predicted Instance Segmentation as a new feature for semantic segmentation. It also…
Token Merging: Your ViT But Faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +3
We introduce Token Merging (ToMe), a simple method to increase the throughput of existing ViT models without needing to train. ToMe gradually combines similar tokens in a transform…
Target Driven Instance Detection
Phil Ammirato, Cheng-Yang Fu, Mykhailo Shvets +2
While state-of-the-art general object detectors are getting better and better, there are not many systems specifically designed to take advantage of the instance detection problem.…
End-to-End Visual Editing with a Generatively Pre-Trained Artist
Andrew Brown, Cheng-Yang Fu, Omkar Parkhi +2
We consider the targeted image editing problem: blending a region in a source image with a driver image that specifies the desired change. Differently from prior works, we solve th…
Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query Localization
Mengmeng Xu, Yanghao Li, Cheng-Yang Fu +3
This paper deals with the problem of localizing objects in image and video datasets from visual exemplars. In particular, we focus on the challenging problem of egocentric visual q…