8 papers
STEAM: Stable Self-Training with Elastic Matching and Adaptive Purification
Shaoxiang Wang, Kejia Zhang, Haiwei Pan +1
Cross-view geo-localization (CVGL) aims to achieve GPS-free localization by matching drone-view images with corresponding satellite-view images. Existing supervised methods rely on…
Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction
Chun-Peng Chang, Shaoxiang Wang, Alain Pagani +2
Modern image encoders achieve high generalization by decoupling semantic meaning from resolution, an ability yet to be fully realized in the 3D domain. We investigate the failure o…
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
Christen Millerdurai, Shaoxiang Wang, Yaxu Xie +3
Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, te…
ReLaGS: Relational Language Gaussian Splatting
Yaxu Xie, Abdalla Arafa, Alireza Javanmardi +5
Achieving unified 3D perception and reasoning across tasks such as segmentation, retrieval, and relation understanding remains challenging, as existing methods are either object-ce…
TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
Alireza Javanmardi, Pragati Jaiswal, Tewodros Amberbir Habtegebrial +4
Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion fr…
Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
Shaoxiang Wang, Shihong Zhang, Christen Millerdurai +3
Despite recent advances in single-object front-facing inpainting using NeRF and 3D Gaussian Splatting (3DGS), inpainting in complex 360° scenes remains largely underexplored. This…