3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Language-Image Alignment with Fixed Text Encoders
Jingfeng Yang, Ziyang Wu, Yue Zhao +1
Currently, the most dominant approach to establishing language-image alignment is to pre-train text and image encoders jointly through contrastive learning, such as CLIP and its va…
cs.CV2024★ 3 cited
Forging Vision Foundation Models for Autonomous Driving: Challenges, Methodologies, and Opportunities
Xu Yan, Haiming Zhang, Yingjie Cai +13
The rise of large foundation models, trained on extensive datasets, is revolutionizing the field of AI. Models such as SAM, DALL-E2, and GPT-4 showcase their adaptability by extrac…
cs.CV2023★ 1 cited
LEAP: Liberate Sparse-view 3D Modeling from Camera Poses
Hanwen Jiang, Zhenyu Jiang, Yue Zhao +1
Are camera poses necessary for multi-view 3D modeling? Existing approaches predominantly assume access to accurate camera poses. While this assumption might hold for dense views, a…