14 citations · 44 across the 6 of their papers we have counts for
6 papers
TransVOS: Video Object Segmentation with Transformers
Jianbiao Mei, Mengmeng Wang, Yeneng Lin +2
Recently, Space-Time Memory Network (STM) based methods have achieved state-of-the-art performance in semi-supervised video object segmentation (VOS). A crucial problem in this tas…
Unsupervised Sound Localization via Iterative Contrastive Learning
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee +2
Sound localization aims to find the source of the audio signal in the visual scene. However, it is labor-intensive to annotate the correlations between the signals sampled from the…
MVHM: A Large-Scale Multi-View Hand Mesh Benchmark for Accurate 3D Hand Pose Estimation
Liangjian Chen, Shih-Yao Lin, Yusheng Xie +2
Estimating 3D hand poses from a single RGB image is challenging because depth ambiguity leads the problem ill-posed. Training hand pose estimators with 3D hand mesh annotations and…
Temporal-Aware Self-Supervised Learning for 3D Hand Pose and Mesh Estimation in Videos
Liangjian Chen, Shih-Yao Lin, Yusheng Xie +2
Estimating 3D hand pose directly from RGB imagesis challenging but has gained steady progress recently bytraining deep models with annotated 3D poses. Howeverannotating 3D poses is…
MM-Hand: 3D-Aware Multi-Modal Guided Hand Generative Network for 3D Hand Pose Synthesis
Zhenyu Wu, Duc Hoang, Shih-Yao Lin +5
Estimating the 3D hand pose from a monocular RGB image is important but challenging. A solution is training on large-scale RGB hand images with accurate 3D hand keypoint annotation…
Every Pixel Matters: Center-aware Feature Alignment for Domain Adaptive Object Detector
Cheng-Chun Hsu, Yi-Hsuan Tsai, Yen-Yu Lin +1
A domain adaptive object detector aims to adapt itself to unseen domains that may contain variations of object appearance, viewpoints or backgrounds. Most existing methods adopt fe…