4 papers
CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template Decomposition
Hongwen Zhang, Siyou Lin, Ruizhi Shao +5
Creating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deform…
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
Ziwen Li, Xin Wang, Hanlue Zhang +8
The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs…
SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation
Han Huang, Han Sun, Ningzhong Liu +2
Driven by autonomous driving's demands for precise 3D perception, 3D semantic occupancy prediction has become a pivotal research topic. Unlike bird's-eye-view (BEV) methods, which…
Depth3DLane: Monocular 3D Lane Detection via Depth Prior Distillation
Dongxin Lyu, Han Huang, Cheng Tan +1
Monocular 3D lane detection is challenging due to the difficulty in capturing depth information from single-camera images. A common strategy involves transforming front-view (FV) i…