5 papers
NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
Dijia Zhan, Jinyi Li, Chenxi Zheng +4
Existing Vision-Language Navigation (VLN) methods typically adopt an egocentric, step-by-step paradigm, which struggles with error accumulation and limits efficiency. While recent…
RecDreamer: Consistent Text-to-3D Generation via Uniform Score Distillation
Chenxi Zheng, Yihong Lin, Bangzhen Liu +3
Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. Thi…
Rotation-Adaptive Point Cloud Domain Generalization via Intricate Orientation Learning
Bangzhen Liu, Chenxi Zheng, Xuemiao Xu +3
The vulnerability of 3D point cloud analysis to unpredictable rotations poses an open yet challenging problem: orientation-aware 3D domain generalization. Cross-domain robustness a…
VrdONE: One-stage Video Visual Relation Detection
Xinjie Jiang, Chenxi Zheng, Xuemiao Xu +4
Video Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyo…
Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
Zikai Huang, Xuemiao Xu, Cheng Xu +4
Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging,…