7k citations
- University of California, Santa BarbaraUS110 papers
- Microsoft Research (United Kingdom)GB57 papers
- ETH ZurichCH48 papers
- Carnegie Mellon UniversityUS45 papers
- University of California, BerkeleyUS45 papers
- University of Maryland, College ParkUS44 papers
- Stanford UniversityUS40 papers
- University of WashingtonUS40 papers
- Princeton UniversityUS35 papers
- Cornell UniversityUS33 papers
- California Institute of TechnologyUS28 papers
- Microsoft Research New York City (United States)27 papers
12 papers · 2 filters
Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers
Yanhong Zeng, Huan Yang, Hongyang Chao +2
We present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a fu…
Bootstrap Your Object Detector via Mixed Training
Mengde Xu, Zheng Zhang, Fangyun Wei +5
We introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by ut…
SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation
Abhinav Moudgil, Arjun Majumdar, Harsh Agrawal +2
Natural language instructions for visual navigation often use scene descriptions (e.g., "bedroom") and object references (e.g., "green chairs") to provide a breadcrumb trail to a g…
Semi-Supervised Semantic Segmentation via Adaptive Equalization Learning
Hanzhe Hu, Fangyun Wei, Han Hu +3
Due to the limited and even imbalanced data, semi-supervised semantic segmentation tends to have poor performance on some certain categories, e.g., tailed categories in Cityscapes…
PatchMatch-RL: Deep MVS with Pixelwise Depth, Normal, and Visibility
Jae Yong Lee, Joseph DeGol, Chuhang Zou +1
Recent learning-based multi-view stereo (MVS) methods show excellent performance with dense cameras and small depth ranges. However, non-learning based approaches still outperform…
PoseRN: A 2D pose refinement network for bias-free multi-view 3D human pose estimation
Akihiko Sayo, Diego Thomas, Hiroshi Kawasaki +2
We propose a new 2D pose refinement network that learns to predict the human bias in the estimated 2D pose. There are biases in 2D pose estimations that are due to differences betw…