14 papers
Monocular Depth Estimation from a Single Image: Progress and Opportunities
Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun +4
Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous dri…
ArtiMo: Agent-Driven Articulated Mesh Animation
Chunyu Zou, Peng Dai, Yi-Hua Huang +4
Animating articulated 3D meshes via text requires satisfying strict kinematic constraints, modeling causal interactions between parts, and achieving instruction fidelity. Due to th…
Geometry-Instructed Video Editing
Chirui Chang, Xiaoyang Lyu, Yi-Hua Huang +7
Object-level geometric edits, including translating, rotating, scaling, duplicating, or removing an object, are routine operations in digital content creation (DCC) workflows, yet…
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
Mingxian Lin, Shengju Qian, Yuqi Liu +9
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per…
Stabilizing Streaming Video Geometry via Dynamic Feature Normalization
Xiaoyang Lyu, Muxin Liu, Xiaoshan Wu +5
Consistent 3D geometry estimation from streaming RGB input is crucial for real-world applications such as autonomous driving, embodied AI, and large-scale reconstruction. While mod…
AniGen: Unified Fields for Animatable 3D Asset Generation
Yi-Hua Huang, Zi-Xin Zou, Yuting He +6
Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodied agents, and animation produc…