collaborators

13 papers

cs.CV2026

MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

Lingyu Kong, Ruicheng Li, Ruicheng Wang +4

Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structu…

cs.CV2026

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

Sicheng Xu, Yu Deng, Shoukang Hu +5

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a…

cs.AI2026

TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning

ZhiYuan Feng, Yu Deng, Ruichuan An +11

In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such re…

cs.CV2026

Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data

Yizhao Xu, Hongyuan Zhu, Caiyun Liu +6

3D editing refers to the ability to apply local or global modifications to 3D assets. Effective 3D editing requires maintaining semantic consistency by performing localized changes…

cs.CV2026

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

Huizhi Liang, Yichao Shen, Yu Deng +5

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D…

cs.CV2026

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

Zhiyuan Feng, Zhaolu Kang, Qijie Wang +16

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent V…