14 citations · 15 across the 4 of their papers we have counts for
6 papers · 1 filter
ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection
Tao Tu, Shun-Po Chuang, Yu-Lun Liu +5
We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods…
ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation
Xuefeng Hu, Ke Zhang, Lu Xia +9
Large-scale Pre-Training Vision-Language Model such as CLIP has demonstrated outstanding performance in zero-shot classification, e.g. achieving 76.3% top-1 accuracy on ImageNet wi…
CC-3DT: Panoramic 3D Object Tracking via Cross-Camera Fusion
Tobias Fischer, Yung-Hsu Yang, Suryansh Kumar +2
To track the 3D locations and trajectories of the other traffic participants at any given time, modern autonomous vehicles are equipped with multiple cameras that cover the vehicle…
360-MLC: Multi-view Layout Consistency for Self-training and Hyper-parameter Tuning
Bolivar Solarte, Chin-Hsuan Wu, Yueh-Cheng Liu +2
We present 360-MLC, a self-training method based on multi-view layout consistency for finetuning monocular room-layout models using unlabeled 360-images only. This can be valuable…
Autoregressive 3D Shape Generation via Canonical Mapping
An-Chieh Cheng, Xueting Li, Sifei Liu +2
With the capacity of modeling long-range dependencies in sequential data, transformers have shown remarkable performances in a variety of generative tasks such as image, audio, and…
Specialize and Fuse: Pyramidal Output Representation for Semantic Segmentation
Chi-Wei Hsiao, Cheng Sun, Hwann-Tzong Chen +1
We present a novel pyramidal output representation to ensure parsimony with our "specialize and fuse" process for semantic segmentation. A pyramidal "output" representation consist…