1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2023
The Role of Linguistic Priors in Measuring Compositional Generalization of Vision-Language Models
Chenwei Wu, Li Erran Li, Stefano Ermon +3
Compositionality is a common property in many modalities including natural languages and images, but the compositional generalization of multi-modal models is not well-understood.…
cs.CV2023★ 1 cited
LiDAR-Based 3D Object Detection via Hybrid 2D Semantic Scene Generation
Haitao Yang, Zaiwei Zhang, Xiangru Huang +5
Bird's-Eye View (BEV) features are popular intermediate scene representations shared by the 3D backbone and the detector head in LiDAR-based object detectors. However, little resea…