7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.CV2024
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
Hao Lu, Xuesong Niu, Jiyao Wang +12
Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos. Despite its success in l…
cs.CV2023★ 7 cited
Fast-BEV: Towards Real-time On-vehicle Bird's-Eye View Perception
Bin Huang, Yangguang Li, Enze Xie +7
Recently, the pure camera-based Bird's-Eye-View (BEV) perception removes expensive Lidar sensors, making it a feasible solution for economical autonomous driving. However, most exi…