16 papers
SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
Anbang Wang, Yuzhuo Ao, Shangzhe Wu +1
Native 3D generative models have achieved remarkable fidelity and speed, yet they suffer from a critical limitation: inability to prescribe precise structural articulations, where…
Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos
Yubo Zhao, Yujin Chai, Yunao Dong +4
Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can…
TransmissiveGS: Residual-Guided Disentangled Gaussian Splatting for Transmissive Scene Reconstruction and Rendering
Zhenyu Liang, Xiao Zhang, Tianchao Li +2
Transmissive scenes are ubiquitous in daily life, yet reconstructing and rendering them remains highly challenging due to the inherent entanglement between near-field reflections f…
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang
Reasoning Video Object Segmentation is a challenging task, aiming at generating a mask sequence from an input video given a complex and implicit text query. While existing works fi…
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
Yuzhuo Ao, Anbang Wang, Yu-Wing Tai +1
Embodied agents often struggle with efficient navigation because they rely primarily on partial egocentric observations, which restrict global foresight and lead to inefficient exp…
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
Shiu-hong Kao, Chak Ho Huang, Huaiqian Liu +2
Existing works of reasoning segmentation often fall short in complex cases, particularly when addressing complicated queries and out-of-domain images. Inspired by the chain-of-thou…