papers

Publications (50)

cs.CV2020

Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation

Bowen Cheng, Maxwell D. Collins, Yukun Zhu +4

In this work, we introduce Panoptic-DeepLab, a simple, strong, and fast system for panoptic segmentation, aiming to establish a solid baseline for bottom-up methods that can achiev…

cs.CV2019

Geo-Aware Networks for Fine-Grained Recognition

Grace Chu, Brian Potetz, Weijun Wang +5

Fine-grained recognition distinguishes among categories with subtle visual differences. In order to differentiate between these challenging visual categories, it is helpful to leve…

cs.CV2025

Epsilon-VAE: Denoising as Visual Decoding

Long Zhao, Sanghyun Woo, Ziyu Wan +6

In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…

cs.CV2020

View-Invariant Probabilistic Embedding for Human Pose

Jennifer J. Sun, Jiaping Zhao, Liang-Chieh Chen +3

Depictions of similar human body configurations can vary with changing viewpoints. Using only 2D information, we would like to enable vision algorithms to recognize similarity in h…

cs.CV2021

Exploring Temporal Granularity in Self-Supervised Video Representation Learning

Rui Qian, Yeqing Li, Liangzhe Yuan +7

This work presents a self-supervised learning framework named TeG to explore Temporal Granularity in learning video representations. In TeG, we sample a long clip from a video and…

cs.CV2021

View-Invariant, Occlusion-Robust Probabilistic Embedding for Human Pose

Ting Liu, Jennifer J. Sun, Long Zhao +6

Recognition of human poses and actions is crucial for autonomous systems to interact smoothly with people. However, cameras generally capture human poses in 2D as images and videos…

cs.CV2021

Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization

Long Zhao, Yuxiao Wang, Jiaping Zhao +7

We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-vie…

cs.CV2020

ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation

Siyuan Qiao, Yukun Zhu, Hartwig Adam +2

In this paper, we present ViP-DeepLab, a unified model attempting to tackle the long-standing and challenging inverse projection problem in vision, which we model as restoring the…

cs.CV2024

VideoGLUE: Video General Understanding Evaluation of Foundation Models

Liangzhe Yuan, Nitesh Bharadwaj Gundavarapu, Long Zhao +14

We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recog…

cs.CV2021

MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers

Huiyu Wang, Yukun Zhu, Hartwig Adam +2

We present MaX-DeepLab, the first end-to-end model for panoptic segmentation. Our approach simplifies the current pipeline that depends heavily on surrogate sub-tasks and hand-desi…

cs.CV2021

EEV: A Large-Scale Dataset for Studying Evoked Expressions from Video

Jennifer J. Sun, Ting Liu, Alan S. Cowen +3

Videos can evoke a range of affective responses in viewers. The ability to predict evoked affect from a video, before viewers watch the video, can help in content creation and vide…

cs.CV2018

NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications

Tien-Ju Yang, Andrew Howard, Bo Chen +5

This work proposes an algorithm, called NetAdapt, that automatically adapts a pre-trained deep neural network to a mobile platform given a resource budget. While many existing algo…

cs.CV2023

Adaptive Transformers for Robust Few-shot Cross-domain Face Anti-spoofing

Hsin-Ping Huang, Deqing Sun, Yaojie Liu +5

While recent face anti-spoofing methods perform well under the intra-domain setups, an effective approach needs to account for much larger appearance variations of images acquired…

cs.CV2021

DeepLab2: A TensorFlow Library for Deep Labeling

Mark Weber, Huiyu Wang, Siyuan Qiao +12

DeepLab2 is a TensorFlow library for deep labeling, aiming to provide a state-of-the-art and easy-to-use TensorFlow codebase for general dense pixel prediction problems in computer…

cs.CV2020

Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset

Menglin Jia, Mengyun Shi, Mikhail Sirotenko +5

In this work we explore the task of instance segmentation with attribute localization, which unifies instance segmentation (detect and segment each object instance) and fine-graine…

cs.CV2017

Rethinking Atrous Convolution for Semantic Image Segmentation

Liang-Chieh Chen, George Papandreou, Florian Schroff +1

In this work, we revisit atrous convolution, a powerful tool to explicitly adjust filter's field-of-view as well as control the resolution of feature responses computed by Deep Con…

cs.CV2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

Dan Kondratyuk, Lijun Yu, Xiuye Gu +28

We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-on…

cs.CV2023

Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception

Hassan Akbari, Dan Kondratyuk, Yin Cui +3

We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, vid…

cs.CV2022

CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation

Qihang Yu, Huiyu Wang, Dahun Kim +6

We propose Clustering Mask Transformer (CMT-DeepLab), a transformer-based framework for panoptic segmentation designed around clustering. It rethinks the existing transformer archi…

cs.LG2022

Surrogate Gap Minimization Improves Sharpness-Aware Training

Juntang Zhuang, Boqing Gong, Liangzhe Yuan +6

The recently proposed Sharpness-Aware Minimization (SAM) improves generalization by minimizing a \textit{perturbed loss} defined as the maximum loss within a neighborhood in the pa…

cs.CV2019

Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation

Chenxi Liu, Liang-Chieh Chen, Florian Schroff +4

Recently, Neural Architecture Search (NAS) has successfully identified neural network architectures that exceed human designed ones on large-scale image classification. In this pap…

cs.CV2024

Video Creation by Demonstration

Yihong Sun, Hao Zhou, Liangzhe Yuan +7

We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physical…

cs.CV2023

Improving Zero-shot Generalization and Robustness of Multi-modal Models

Yunhao Ge, Jie Ren, Andrew Gallagher +6

Multi-modal image-text models such as CLIP and LiT have demonstrated impressive performance on image classification benchmarks and their zero-shot generalization ability is particu…

cs.CV2018

InclusiveFaceNet: Improving Face Attribute Detection with Race and Gender Diversity

Hee Jung Ryu, Hartwig Adam, Margaret Mitchell

We demonstrate an approach to face attribute detection that retains or improves attribute detection accuracy across gender and race subgroups by learning demographic information pr…

cs.CV2018

The iNaturalist Species Classification and Detection Dataset

Grant Van Horn, Oisin Mac Aodha, Yang Song +6

Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imb…

cs.CV2019

The iMaterialist Fashion Attribute Dataset

Sheng Guo, Weilin Huang, Xiao Zhang +6

Large-scale image databases such as ImageNet have significantly advanced image classification and other visual recognition tasks. However much of these datasets are constructed onl…

cs.CV2025

VideoPrism: A Foundational Visual Encoder for Video Understanding

Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16

We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus…

cs.CV2024

Distilling Vision-Language Models on Millions of Videos

Yue Zhao, Long Zhao, Xingyi Zhou +9

The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply…

cs.CV2023

MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models

Chenglin Yang, Siyuan Qiao, Qihang Yu +5

This paper presents MOAT, a family of neural networks that build on top of MObile convolution (i.e., inverted residual blocks) and ATtention. Unlike the current works that stack se…

cs.CV2019

Panoptic-DeepLab

Bowen Cheng, Maxwell D. Collins, Yukun Zhu +4

We present Panoptic-DeepLab, a bottom-up and single-shot approach for panoptic segmentation. Our Panoptic-DeepLab is conceptually simple and delivers state-of-the-art results. In p…

cs.CV2020

Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation

Liang-Chieh Chen, Raphael Gontijo Lopes, Bowen Cheng +5

Supervised learning in large discriminative models is a mainstay for modern computer vision. Such an approach necessitates investing in large-scale human-annotated datasets for ach…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CV2021

STEP: Segmenting and Tracking Every Pixel

Mark Weber, Jun Xie, Maxwell Collins +10

The task of assigning semantic classes and track identities to every pixel in a video is called video panoptic segmentation. Our work is the first that targets this task in a real-…

cs.CV2021

2.5D Visual Relationship Detection

Yu-Chuan Su, Soravit Changpinyo, Xiangning Chen +8

Visual 2.5D perception involves understanding the semantics and geometry of a scene through reasoning about object relationships with respect to the viewer in an environment. Howev…

cs.CV2019

FEELVOS: Fast End-to-End Embedding Learning for Video Object Segmentation

Paul Voigtlaender, Yuning Chai, Florian Schroff +3

Many of the recent successful methods for video object segmentation (VOS) are overly complicated, heavily rely on fine-tuning on the first frame, and/or are slow, and are hence of…

cs.CV2024

SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset

Sagar M. Waghmare, Kimberly Wilber, Dave Hawkey +9

Vision is essential for human navigation. The World Health Organization (WHO) estimates that 43.3 million people were blind in 2020, and this number is projected to reach 61 millio…

cs.CV2017

Learning Unified Embedding for Apparel Recognition

Yang Song, Yuan Li, Bo Wu +3

In apparel recognition, specialized models (e.g. models trained for a particular vertical like dresses) can significantly outperform general models (i.e. models that cover a wide r…

cs.CV2017

MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features

Liang-Chieh Chen, Alexander Hermans, George Papandreou +3

In this work, we tackle the problem of instance segmentation, the task of simultaneously solving object detection and semantic segmentation. Towards this goal, we present a model,…

cs.CV2019

Searching for MobileNetV3

Andrew Howard, Mark Sandler, Grace Chu +9

We present the next generation of MobileNets based on a combination of complementary search techniques as well as a novel architecture design. MobileNetV3 is tuned to mobile phone…

cs.CV2023

PolyMaX: General Dense Prediction with Mask Transformer

Xuan Yang, Liangzhe Yuan, Kimberly Wilber +8

Dense prediction tasks, such as semantic segmentation, depth estimation, and surface normal prediction, can be easily formulated as per-pixel classification (discrete outputs) or r…

cs.CV2023

kMaX-DeepLab: k-means Mask Transformer

Qihang Yu, Huiyu Wang, Siyuan Qiao +5

The rise of transformers in vision tasks not only advances network backbone designs, but also starts a brand-new page to achieve end-to-end image recognition (e.g., object detectio…

cs.CV2022

Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset

Grant Van Horn, Rui Qian, Kimberly Wilber +3

We present a new benchmark dataset, Sapsucker Woods 60 (SSW60), for advancing research on audiovisual fine-grained categorization. While our community has made great strides in fin…

cs.CV2017

MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Andrew G. Howard, Menglong Zhu, Bo Chen +5

We present a class of efficient models called MobileNets for mobile and embedded vision applications. MobileNets are based on a streamlined architecture that uses depth-wise separa…

cs.CV2023

TubeFormer-DeepLab: Video Mask Transformer

Dahun Kim, Jun Xie, Huiyu Wang +6

We present TubeFormer-DeepLab, the first attempt to tackle multiple core video segmentation tasks in a unified manner. Different video segmentation tasks (e.g., video semantic/inst…

cs.CV2018

Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation

Liang-Chieh Chen, Yukun Zhu, George Papandreou +2

Spatial pyramid pooling module or encode-decoder structure are used in deep neural networks for semantic segmentation task. The former networks are able to encode multi-scale conte…

cs.LG2017

Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Benoit Jacob, Skirmantas Kligys, Bo Chen +5

The rising popularity of intelligent mobile devices and the daunting computational cost of deep learning-based models call for efficient and accurate on-device inference schemes. W…

cs.CV2023

Unified Visual Relationship Detection with Vision and Language Models

Long Zhao, Liangzhe Yuan, Boqing Gong +5

This work focuses on training a single visual relationship detector predicting over the union of label spaces from multiple datasets. Merging labels spanning different datasets cou…

cs.CV2018

Searching for Efficient Multi-Scale Architectures for Dense Image Prediction

Liang-Chieh Chen, Maxwell D. Collins, Yukun Zhu +5

The design of neural network architectures is an important component for achieving state-of-the-art performance with machine learning systems across a broad array of tasks. Much wo…

cs.CV2022

Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision

Liangzhe Yuan, Rui Qian, Yin Cui +5

Modern self-supervised learning algorithms typically enforce persistency of instance representations across views. While being very effective on learning holistic image and video r…

cs.CV2020

Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation

Huiyu Wang, Yukun Zhu, Bradley Green +3

Convolution exploits locality for efficiency at a cost of missing long range context. Self-attention has been adopted to augment CNNs with non-local interactions. Recent works prov…