most citedCLIP: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

3 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2024

GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data

Haoyuan Li, Yanpeng Zhou, Yihan Zeng +2

3D Shape represented as point cloud has achieve advancements in multimodal pre-training to align image and language descriptions, which is curial to object identification, classifi…

cs.CV2023

PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion

Guansong Lu, Yuanfan Guo, Jianhua Han +7

Current large-scale diffusion models represent a giant leap forward in conditional image synthesis, capable of interpreting diverse cues like text, human poses, and edges. However,…

cs.CV2023

Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images

Cuican Yu, Guansong Lu, Yihan Zeng +7

Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional…

cs.CV2023

DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination Capability

Runhui Huang, Jianhua Han, Guansong Lu +4

Recently, large-scale diffusion models, e.g., Stable diffusion and DallE2, have shown remarkable results on image synthesis. On the other hand, large-scale cross-modal pre-trained…

cs.CV20233 cited

CLIP: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

Yihan Zeng, Chenhan Jiang, Jiageng Mao +7

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. Howeve…