15 papers
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang +6
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…
: 3D Reconstruction via Relative Regression
Congrong Xu, Huachen Gao, Xingyu Chen +3
Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typical…
Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences
Bao Li, Yuliang Xiu, Zhen Liu
Large-scale text-to-image foundation models have achieved remarkable visual realism, yet generating human images with correct anatomical structures remains challenging. Existing ap…
ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Datasets
Xiaoben Li, Jingyi Wu, Zeyu Cai +3
Human body fitting, which aligns parametric body models such as SMPL to raw 3D point clouds of clothed humans, serves as a crucial first step for downstream tasks like animation an…
OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction
Zeyu Cai, Yuliang Xiu, Renke Wang +8
Fitting an underlying body model to 3D clothed human assets has been extensively studied, yet most approaches focus on either single-modal inputs such as point clouds or multi-view…
GaussiAnimate: Reconstruct and Rig Animatable Categories with Level of Dynamics
Jiaxin Wang, Dongxin Lyu, Zeyu Cai +4
Free-form bones, that conform closely to the surface, can effectively capture non-rigid deformations, but lack a kinematic structure necessary for intuitive control. Thus, we propo…