activity
20242026
collaborators
Showing 2024 · cs.CVShow all

5 papers · 2 filters

cs.CV2024

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

Hao Wu, Zhihang Zhong, Xiao Sun

Image captioning models often suffer from performance degradation when applied to novel datasets, as they are typically trained on domain-specific data. To enhance generalization i…

cs.CV2024

MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks

Yifei Liu, Zhihang Zhong, Yifan Zhan +2

While 3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and real-time rendering, the high memory consumption due to the use of millions o…

cs.CV2024

X as Supervision: Contending with Depth Ambiguity in Unsupervised Monocular 3D Pose Estimation

Yuchen Yang, Xuanyi Liu, Xing Gao +2

Recent unsupervised methods for monocular 3D pose estimation have endeavored to reduce dependence on limited annotated 3D data, but most are solely formulated in 2D space, overlook…

cs.CV2024

Sequential Gaussian Avatars with Hierarchical Motion Context

Wangze Xu, Yifan Zhan, Zhihang Zhong +1

The emergence of neural rendering has significantly advanced the rendering quality of 3D human avatars, with the recently popular 3DGS technique enabling real-time performance. How…

cs.CV2024

Motion-Aware Animatable Gaussian Avatars Deblurring

Muyao Niu, Yifan Zhan, Qingtian Zhu +5

The creation of 3D human avatars from multi-view videos is a significant yet challenging task in computer vision. However, existing techniques rely on high-quality, sharp images as…