Publications (129)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
Haoxuan You, Rui Sun, Zhecan Wang +5
The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot…
Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
Yu-Gang Jiang, Zuxuan Wu, Jun Wang +2
In this paper, we study the challenging problem of categorizing videos according to high-level semantics such as the existence of a particular human action or a complex event. Alth…
CamSwarm: Instantaneous Smartphone Camera Arrays for Collaborative Photography
Yan Wang, Jue Wang, Shih-Fu Chang
Camera arrays (CamArrays) are widely used in commercial filming projects for achieving special visual effects such as bullet time effect, but are very expensive to set up. We propo…
Learning from Children: Improving Image-Caption Pretraining via Curriculum
Hammad A. Ayyubi, Rahul Lokesh, Alireza Zareian +2
Image-caption pretraining has been quite successfully used for downstream vision tasks like zero-shot image classification and object detection. However, image-caption pretraining…
Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
Liunian Harold Li, Haoxuan You, Zhecan Wang +3
Pre-trained contextual vision-and-language (V&L) models have achieved impressive performance on various benchmarks. However, existing models require a large amount of parallel imag…
PatternNet: Visual Pattern Mining with Deep Neural Network
Hongzhi Li, Joseph G. Ellis, Lei Zhang +1
Visual patterns represent the discernible regularity in the visual world. They capture the essential nature of visual objects or scenes. Understanding and modeling visual patterns…
Context-Gated Convolution
Xudong Lin, Lin Ma, Wei Liu +1
As the basic building block of Convolutional Neural Networks (CNNs), the convolutional layer is designed to extract local patterns and lacks the ability to model global context in…
DeepSentiBank: Visual Sentiment Concept Classification with Deep Convolutional Neural Networks
Tao Chen, Damian Borth, Trevor Darrell +1
This paper introduces a visual sentiment concept classification method based on deep convolutional neural networks (CNNs). The visual sentiment concepts are adjective noun pairs (A…
Multimodal Social Media Analysis for Gang Violence Prevention
Philipp Blandfort, Desmond Patton, William R. Frey +9
Gang violence is a severe issue in major cities across the U.S. and recent studies [Patton et al. 2017] have found evidence of social media communications that can be linked to suc…
SVM for learning with label proportions
Felix X. Yu, Dong Liu, Sanjiv Kumar +2
We study the problem of learning with label proportions in which the training data is provided in groups and only the proportion of each class in each group is known. We propose a…
On the Difficulty of Nearest Neighbor Search
Junfeng He, Sanjiv Kumar, Shih-Fu Chang
Fast approximate nearest neighbor (NN) search in large databases is becoming popular. Several powerful learning-based formulations have been proposed recently. However, not much at…
Video Event Extraction via Tracking Visual States of Arguments
Guang Yang, Manling Li, Jiajie Zhang +3
Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the…
Compact Hyperplane Hashing with Bilinear Functions
Wei Liu, Jun Wang, Yadong Mu +2
Hyperplane hashing aims at rapidly searching nearest points to a hyperplane, and has shown practical impact in scaling up active learning with SVMs. Unfortunately, the existing ran…
WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
Jiawei Ma, Yulei Niu, Shiyuan Huang +2
Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mo…
Learning To Recognize Procedural Activities with Distant Supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius +3
In this paper we consider the problem of classifying fine-grained, multi-step activities (e.g., cooking different recipes, making disparate home improvements, creating various form…
Learning to Learn Words from Visual Scenes
DÃdac SurÃs, Dave Epstein, Heng Ji +2
Language acquisition is the process of learning words from the surrounding scene. We introduce a meta-learning framework that learns how to learn word representations from unconstr…
Enhanced Chart Understanding in Vision and Language Task via Cross-modal Pre-training on Plot Table Pairs
Mingyang Zhou, Yi R. Fung, Long Chen +3
Building cross-model intelligence that can understand charts and communicate the salient information hidden behind them is an appealing challenge in the vision and language(V+L) co…
MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and Grounding
Revanth Gangi Reddy, Xilin Rui, Manling Li +9
Recently, there has been an increasing interest in building question answering (QA) models that reason across multiple modalities, such as text and images. However, QA using images…
Fine-Grained Visual Entailment
Christopher Thomas, Yipeng Zhang, Shih-Fu Chang
Visual entailment is a recently proposed multimodal reasoning task where the goal is to predict the logical relationship of a piece of text to an image. In this paper, we propose a…
Deep Learning Guided Building Reconstruction from Satellite Imagery-derived Point Clouds
Bo Xu, Xu Zhang, Zhixin Li +3
3D urban reconstruction of buildings from remotely sensed imagery has drawn significant attention during the past two decades. While aerial imagery and LiDAR provide higher resolut…
Columbia MVSO Image Sentiment Dataset
Vaidehi Dalmia, Hongyi Liu, Shih-Fu Chang
The Multilingual Visual Sentiment Ontology (MVSO) consists of 15,600 concepts in 12 different languages that are strongly related to emotions and sentiments expressed in images. Th…
Counterfactual Critic Multi-Agent Training for Scene Graph Generation
Long Chen, Hanwang Zhang, Jun Xiao +3
Scene graphs -- objects as nodes and visual relationships as edges -- describe the whereabouts and interactions of the things and stuff in an image for comprehensive scene understa…
Weakly-Supervised Temporal Article Grounding
Long Chen, Yulei Niu, Brian Chen +6
Given a long untrimmed video and natural language queries, video grounding (VG) aims to temporally localize the semantically-aligned video segments. Almost all existing VG work hol…
Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
Zhenhailong Wang, Manling Li, Ruochen Xu +10
The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples, such as domain-specific captioning, question…
Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
Haoxuan You, Luowei Zhou, Bin Xiao +5
Large-scale multi-modal contrastive pre-training has demonstrated great utility to learn transferable features for a range of downstream tasks by mapping multiple modalities into a…
COVID-19 Literature Knowledge Graph Construction and Drug Repurposing Report Generation
Qingyun Wang, Manling Li, Xuan Wang +24
To combat COVID-19, both clinicians and scientists need to digest vast amounts of relevant biomedical knowledge in scientific literature to understand the disease mechanism and rel…
Generic Instance Search and Re-identification from One Example via Attributes and Categories
Ran Tao, Arnold W. M. Smeulders, Shih-Fu Chang
This paper aims for generic instance search from one example where the instance can be an arbitrary object like shoes, not just near-planar and one-sided instances like buildings a…
Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
Victor Campos, Brendan Jou, Xavier Giro-i-Nieto +2
Recurrent Neural Networks (RNNs) continue to show outstanding performance in sequence modeling tasks. However, training RNNs on long sequences often face challenges like slow infer…
Distributed Low-rank Subspace Segmentation
Ameet Talwalkar, Lester Mackey, Yadong Mu +2
Vision problems ranging from image clustering to motion segmentation to semi-supervised learning can naturally be framed as subspace segmentation problems, in which one aims to rec…
Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond
Zhecan Wang, Long Chen, Haoxuan You +6
Vision-language (VL) understanding tasks evaluate models' comprehension of complex visual scenes through multiple-choice questions. However, we have identified two dataset biases t…
Bridging Knowledge Graphs to Generate Scene Graphs
Alireza Zareian, Svebor Karaman, Shih-Fu Chang
Scene graphs are powerful representations that parse images into their abstract semantic elements, i.e., objects and their interactions, which facilitates visual comprehension and…
Find Someone Who: Visual Commonsense Understanding in Human-Centric Grounding
Haoxuan You, Rui Sun, Zhecan Wang +2
From a visual scene containing multiple people, human is able to distinguish each individual given the context descriptions about what happened before, their mental/physical states…
Unsupervised Rank-Preserving Hashing for Large-Scale Image Retrieval
Svebor Karaman, Xudong Lin, Xuefeng Hu +1
We propose an unsupervised hashing method which aims to produce binary codes that preserve the ranking induced by a real-valued representation. Such compact hash codes enable the c…
Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning
Kung-Hsiang Huang, Mingyang Zhou, Hou Pong Chan +5
Recent advancements in large vision-language models (LVLMs) have led to significant progress in generating natural language descriptions for visual content and thus enhancing vario…
Partner-Assisted Learning for Few-Shot Image Classification
Jiawei Ma, Hanchen Xie, Guangxing Han +3
Few-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learn…
Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps
Shih-Fu Chang, Alex Hauptmann, Louis-Philippe Morency +18
With the transformative technologies and the rapidly changing global R&D landscape, the multimedia and multimodal community is now faced with many new opportunities and uncertainti…
CLIP-Event: Connecting Text and Images with Event Structures
Manling Li, Ruochen Xu, Shuohang Wang +6
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing v…
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
Kung-Hsiang Huang, Hou Pong Chan, Yi R. Fung +5
Data visualization in the form of charts plays a pivotal role in data analysis, offering critical insights and aiding in informed decision-making. Automatic chart understanding has…
Multi-granularity Generator for Temporal Action Proposal
Yuan Liu, Lin Ma, Yifeng Zhang +2
Temporal action proposal generation is an important task, aiming to localize the video segments containing human actions in an untrimmed video. In this paper, we propose a multi-gr…
Unsupervised Embedding Learning via Invariant and Spreading Instance Feature
Mang Ye, Xu Zhang, Pong C. Yuen +1
This paper studies the unsupervised embedding learning problem, which requires an effective similarity measurement between samples in low-dimensional embedding space. Motivated by…
Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks
Long Chen, Hanwang Zhang, Jun Xiao +2
We propose a novel framework called Semantics-Preserving Adversarial Embedding Network (SP-AEN) for zero-shot visual recognition (ZSL), where test images and their classes are both…
Supervised Masked Knowledge Distillation for Few-Shot Transformers
Han Lin, Guangxing Han, Jiawei Ma +3
Vision Transformers (ViTs) emerge to achieve impressive performance on many data-abundant computer vision tasks by capturing long-range dependencies among local features. However,…
Low-shot Learning via Covariance-Preserving Adversarial Augmentation Networks
Hang Gao, Zheng Shou, Alireza Zareian +2
Deep neural networks suffer from over-fitting and catastrophic forgetting when trained with small data. One natural remedy for this problem is data augmentation, which has been rec…
Beyond Grounding: Extracting Fine-Grained Event Hierarchies Across Modalities
Hammad A. Ayyubi, Christopher Thomas, Lovish Chum +8
Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of c…
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
Hammad Ayyubi, Xuande Feng, Junzhang Liu +3
The task of predicting time and location from images is challenging and requires complex human-like puzzle-solving ability over different clues. In this work, we formalize this abi…
Towards Train-Test Consistency for Semi-supervised Temporal Action Localization
Xudong Lin, Zheng Shou, Shih-Fu Chang
Recently, Weakly-supervised Temporal Action Localization (WTAL) has been densely studied but there is still a large gap between weakly-supervised models and fully-supervised models…
Video in 10 Bits: Few-Bit VideoQA for Efficiency and Privacy
Shiyuan Huang, Robinson Piramuthu, Shih-Fu Chang +1
In Video Question Answering (VideoQA), answering general questions about a video requires its visual information. Yet, video often contains redundant information irrelevant to the…
Variational Context: Exploiting Visual and Textual Context for Grounding Referring Expressions
Yulei Niu, Hanwang Zhang, Zhiwu Lu +1
We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., ``largest elephant standing behind baby elephant''. This is a general yet challenging vis…
Unifying Specialist Image Embedding into Universal Image Embedding
Yang Feng, Futang Peng, Xu Zhang +7
Deep image embedding provides a way to measure the semantic similarity of two images. It plays a central role in many applications such as image search, face verification, and zero…
AutoLoc: Weakly-supervised Temporal Action Localization
Zheng Shou, Hang Gao, Lei Zhang +2
Temporal Action Localization (TAL) in untrimmed video is important for many applications. But it is very expensive to annotate the segment-level ground truth (action class and temp…
Non-Sequential Graph Script Induction via Multimedia Grounding
Yu Zhou, Sha Li, Manling Li +4
Online resources such as WikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. However, the scripts…
Grounding Referring Expressions in Images by Variational Context
Hanwang Zhang, Yulei Niu, Shih-Fu Chang
We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., "largest elephant standing behind baby elephant". This is a general yet challenging visio…
DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
Zheng Shou, Xudong Lin, Yannis Kalantidis +4
Motion has shown to be useful for video understanding, where motion is typically represented by optical flow. However, computing flow from video frames is very time-consuming. Rece…
Multi-Modal Multi-Scale Deep Learning for Large-Scale Image Annotation
Yulei Niu, Zhiwu Lu, Ji-Rong Wen +2
Image annotation aims to annotate a given image with a variable number of class labels corresponding to diverse visual concepts. In this paper, we address two main issues in large-…
General Partial Label Learning via Dual Bipartite Graph Autoencoder
Brian Chen, Bo Wu, Alireza Zareian +2
We formulate a practical yet challenging problem: General Partial Label Learning (GPLL). Compared to the traditional Partial Label Learning (PLL) problem, GPLL relaxes the supervis…
Joint Multimedia Event Extraction from Video and Article
Brian Chen, Xudong Lin, Christopher Thomas +5
Visual and textual modalities contribute complementary information about events described in multimedia documents. Videos contain rich dynamics and detailed unfoldings of events, w…
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
Rui Sun, Zhecan Wang, Haoxuan You +3
Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model's reasoning ability to understand the semantics of the visual world and natural…
Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment
Guangxing Han, Shiyuan Huang, Jiawei Ma +2
Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object…
Few-Shot Object Detection with Fully Cross-Transformer
Guangxing Han, Jiawei Ma, Shiyuan Huang +2
Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-lea…
Circulant Binary Embedding
Felix X. Yu, Sanjiv Kumar, Yunchao Gong +1
Binary embedding of high-dimensional data requires long codes to preserve the discriminative power of the input space. Traditional binary coding methods often suffer from very high…
Detecting and Simulating Artifacts in GAN Fake Images
Xu Zhang, Svebor Karaman, Shih-Fu Chang
To detect GAN generated images, conventional supervised machine learning algorithms require collection of a number of real and fake images from the targeted GAN model. However, the…
Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos
Brian Chen, Andrew Rouditchenko, Kevin Duarte +10
Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data…
SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense Reasoning
Zhecan Wang, Haoxuan You, Liunian Harold Li +5
Answering complex questions about images is an ambitious goal for machine intelligence, which requires a joint understanding of images, text, and commonsense knowledge, as well as…
ConvNet Architecture Search for Spatiotemporal Feature Learning
Du Tran, Jamie Ray, Zheng Shou +2
Learning image representations with ConvNets by pre-training on ImageNet has proven useful across many visual understanding tasks including object detection, semantic segmentation,…
Model-Driven Feed-Forward Prediction for Manipulation of Deformable Objects
Yinxiao Li, Yan Wang, Yonghao Yue +5
Robotic manipulation of deformable objects is a difficult problem especially because of the complexity of the many different ways an object can deform. Searching such a high dimens…
DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection
Jiawei Ma, Yulei Niu, Jincheng Xu +3
Generalized few-shot object detection aims to achieve precise detection on both base classes with abundant annotations and novel classes with limited training data. Existing approa…
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
Hassan Akbari, Liangzhe Yuan, Rui Qian +4
We present a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures. Specifically, our Video-Audio-Text Transformer…
Ferret: Refer and Ground Anything Anywhere at Any Granularity
Haoxuan You, Haotian Zhang, Zhe Gan +6
We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding op…
Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding
Long Chen, Wenbo Ma, Jun Xiao +2
The prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to…
Learning Visual Commonsense for Robust Scene Graph Generation
Alireza Zareian, Zhecan Wang, Haoxuan You +1
Scene graph generation models understand the scene through object and predicate recognition, but are prone to mistakes due to the challenges of perception in the wild. Perception e…
CDSA: Cross-Dimensional Self-Attention for Multivariate, Geo-tagged Time Series Imputation
Jiawei Ma, Zheng Shou, Alireza Zareian +3
Many real-world applications involve multivariate, geo-tagged time series data: at each location, multiple sensors record corresponding measurements. For example, air quality monit…
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Haotian Zhang, Haoxuan You, Philipp Dufter +8
While Ferret seamlessly integrates regional understanding into the Large Language Model (LLM) to facilitate its referring and grounding capability, it poses certain limitations: co…
VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
Xudong Lin, Gedas Bertasius, Jue Wang +3
We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, whic…
Training with Streaming Annotation
Tongtao Zhang, Heng Ji, Shih-Fu Chang +1
In this paper, we address a practical scenario where training data is released in a sequence of small-scale batches and annotation in earlier phases has lower quality than the late…
Deep Cross Residual Learning for Multitask Visual Recognition
Brendan Jou, Shih-Fu Chang
Residual learning has recently surfaced as an effective means of constructing very deep neural networks for object recognition. However, current incarnations of residual networks d…
Building A Large Concept Bank for Representing Events in Video
Yin Cui, Dong Liu, Jiawei Chen +1
Concept-based video representation has proven to be effective in complex event detection. However, existing methods either manually design concepts or directly adopt concept librar…
Visual Affect Around the World: A Large-scale Multilingual Visual Sentiment Ontology
Brendan Jou, Tao Chen, Nikolaos Pappas +3
Every culture and language is unique. Our work expressly focuses on the uniqueness of culture and language in relation to human affect, specifically sentiment and emotion semantics…
Language Models are Causal Knowledge Extractors for Zero-shot Video Question Answering
Hung-Ting Su, Yulei Niu, Xudong Lin +2
Causal Video Question Answering (CVidQA) queries not only association or temporal relations but also causal relations in a video. Existing question synthesis methods pre-trained qu…
On Binary Embedding using Circulant Matrices
Felix X. Yu, Aditya Bhaskara, Sanjiv Kumar +2
Binary embeddings provide efficient and powerful ways to perform operations on large scale data. However binary embedding typically requires long codes in order to preserve the dis…
Localizing Actions from Video Labels and Pseudo-Annotations
Pascal Mettes, Cees G. M. Snoek, Shih-Fu Chang
The goal of this paper is to determine the spatio-temporal location of actions in video. Where training from hard to obtain box annotations is the norm, we propose an intuitive and…
Weakly Supervised Visual Semantic Parsing
Alireza Zareian, Svebor Karaman, Shih-Fu Chang
Scene Graph Generation (SGG) aims to extract entities, predicates and their semantic structure from images, enabling deep understanding of visual content, with many applications su…
Deep Image Set Hashing
Jie Feng, Svebor Karaman, I-Hong Jhuo +1
In applications involving matching of image sets, the information from multiple images must be effectively exploited to represent each set. State-of-the-art methods use probabilist…
Visual Translation Embedding Network for Visual Relation Detection
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang +1
Visual relations, such as "person ride bike" and "bike next to car", offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting…
Event Specific Multimodal Pattern Mining with Image-Caption Pairs
Hongzhi Li, Joseph G. Ellis, Shih-Fu Chang
In this paper we describe a novel framework and algorithms for discovering image patch patterns from a large corpus of weakly supervised image-caption pairs generated from news eve…
Learning Spread-out Local Feature Descriptors
Xu Zhang, Felix X. Yu, Sanjiv Kumar +1
We propose a simple, yet powerful regularization technique that can be used to significantly improve both the pairwise and triplet losses in learning local feature descriptors. The…
Learning to Hash for Indexing Big Data - A Survey
Jun Wang, Wei Liu, Sanjiv Kumar +1
The explosive growth in big data has attracted much attention in designing efficient indexing and search methods recently. In many critical applications such as large-scale search…
Heated-Up Softmax Embedding
Xu Zhang, Felix Xinnan Yu, Svebor Karaman +2
Metric learning aims at learning a distance which is consistent with the semantic meaning of the samples. The problem is generally solved by learning an embedding for each sample s…
Compact Nonlinear Maps and Circulant Extensions
Felix X. Yu, Sanjiv Kumar, Henry Rowley +1
Kernel approximation via nonlinear random feature maps is widely used in speeding up kernel machines. There are two main challenges for the conventional kernel approximation method…
Cross-media Structured Common Space for Multimedia Event Extraction
Manling Li, Alireza Zareian, Qi Zeng +4
We introduce a new task, MultiMedia Event Extraction (M2E2), which aims to extract events and their arguments from multimedia documents. We develop the first benchmark and collect…
MoDE: CLIP Data Experts via Clustering
Jiawei Ma, Po-Yao Huang, Saining Xie +5
The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web-crawled data. We…
Entity-aware Image Caption Generation
Di Lu, Spencer Whitehead, Lifu Huang +2
Current image captioning approaches generate descriptions which lack specific information, such as named entities that are involved in the images. In this paper we propose a new ta…
Deep Transfer Network: Unsupervised Domain Adaptation
Xu Zhang, Felix Xinnan Yu, Shih-Fu Chang +1
Domain adaptation aims at training a classifier in one dataset and applying it to a related but not identical dataset. One successfully used framework of domain adaptation is to le…
An exploration of parameter redundancy in deep networks with circulant projections
Yu Cheng, Felix X. Yu, Rogerio S. Feris +3
We explore the redundancy of parameters in deep neural networks by replacing the conventional linear projection in fully-connected layers with the circulant projection. The circula…
Online Detection of Action Start in Untrimmed, Streaming Videos
Zheng Shou, Junting Pan, Jonathan Chan +5
We aim to tackle a novel task in action detection - Online Detection of Action Start (ODAS) in untrimmed, streaming videos. The goal of ODAS is to detect the start of an action ins…
Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense
Zhecan Wang, Haoxuan You, Yicheng He +3
Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve compr…
In Defense of Structural Symbolic Representation for Video Event-Relation Prediction
Andrew Lu, Xudong Lin, Yulei Niu +1
Understanding event relationships in videos requires a model to understand the underlying structures of events (i.e. the event type, the associated argument roles, and correspondin…
Analogical Reasoning for Visually Grounded Language Acquisition
Bo Wu, Haoyu Qin, Alireza Zareian +2
Children acquire language subconsciously by observing the surrounding world and listening to descriptions. They can discover the meaning of words even without explicit language kno…
Flow-Distilled IP Two-Stream Networks for Compressed Video Action Recognition
Shiyuan Huang, Xudong Lin, Svebor Karaman +1
Two-stream networks have achieved great success in video recognition. A two-stream network combines a spatial stream of RGB frames and a temporal stream of Optical Flow to make pre…
Beyond Triplet Loss: Meta Prototypical N-tuple Loss for Person Re-identification
Zhizheng Zhang, Cuiling Lan, Wenjun Zeng +2
Person Re-identification (ReID) aims at matching a person of interest across images. In convolutional neural network (CNN) based approaches, loss design plays a vital role in pulli…
Open-Vocabulary Object Detection Using Captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu +1
Despite the remarkable accuracy of deep neural networks in object detection, they are costly to train and scale due to supervision requirements. Particularly, learning more object…