2 citations · 2 across the 10 of their papers we have counts for
13 papers · 1 filter
DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
Qian Wang, Zhenyu Li, Abdelrahman Eldesokey +1
Subject-driven image generation faces an "Identity-Diversity Paradox", where strong identity preservation often leads to rigid and low-diversity outputs. We propose a post-training…
Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning
Xuyang Wang, Zhenyu Li, Jian Ding +4
Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry recons…
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
Zhenyu Li, Sai Kumar Dwivedi, Filip Maric +11
Egocentric human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and…
Any Resolution Any Geometry: From Multi-View To Multi-Patch
Wenqing Cui, Zhenyu Li, Mykola Lavreniuk +4
Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high-resolution prediction remains difficult due to the trade-off between preser…
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
Jian Shi, Michael Birsak, Wenqing Cui +2
This paper revisits the role of positional embeddings (PEs) within vision transformers (ViTs) from a geometric perspective. We show that PEs are not mere token indices but effectiv…
Depth Anything 3: Recovering the Visual Space from Any Views
Haotong Lin, Sili Chen, Junhao Liew +5
We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of…