320 citations · 848 across the 30 of their papers we have counts for
28 papers · 1 filter
TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
Xiyang Dai, Bharat Singh, Joe Yue-Hei Ng +1
We present Temporal Aggregation Network (TAN) which decomposes 3D convolutions into spatial and temporal aggregation blocks. By stacking spatial and temporal convolutions repeatedl…
FA-RPN: Floating Region Proposals for Face Detection
Mahyar Najibi, Bharat Singh, Larry S. Davis
We propose a novel approach for generating region proposals for performing face-detection. Instead of classifying anchor boxes using features from a pixel in the convolutional feat…
Deep Residual Learning in the JPEG Transform Domain
Max Ehrlich, Larry Davis
We introduce a general method of performing Residual Network inference and learning in the JPEG transform domain that allows the network to consume compressed images as input. Our…
AutoFocus: Efficient Multi-Scale Inference
Mahyar Najibi, Bharat Singh, Larry S. Davis
This paper describes AutoFocus, an efficient multi-scale inference algorithm for deep-learning based object detectors. Instead of processing an entire image pyramid, AutoFocus adop…
MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment
Da Zhang, Xiyang Dai, Xin Wang +2
This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of intere…
Explicit Bias Discovery in Visual Question Answering Models
Varun Manjunatha, Nirat Saini, Larry S. Davis
Researchers have observed that Visual Question Answering (VQA) models tend to answer questions by learning statistical biases in the data. For example, their answer to the question…