4 papers
Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Chaolong Yang, Kai Yao, Yuyao Yan +7
Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized…
Towards Training-Free Open-World Classification with 3D Generative Models
Xinzhe Xia, Weiguang Zhao, Yuyao Yan +4
3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring both open-category and open-pose recognition. To addres…
3D-CDRGP: Towards Cross-Device Robotic Grasping Policy in 3D Open World
Weiguang Zhao, Chenru Jiang, Chengrui Zhang +4
Given the diversity of devices and the product upgrades, cross-device research has become an urgent issue that needs to be tackled. To this end, we pioneer in probing the cross-dev…
SSD: Towards Better Text-Image Consistency Metric in Text-to-Image Generation
Zhaorui Tan, Xi Yang, Zihan Ye +4
Generating consistent and high-quality images from given texts is essential for visual-language understanding. Although impressive results have been achieved in generating high-qua…