1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Clapper: Compact Learning and Video Representation in VLMs
Lingyu Kong, Hongzhi Zhang, Jingyuan Zhang +4
Current vision-language models (VLMs) have demonstrated remarkable capabilities across diverse video understanding applications. Designing VLMs for video inputs requires effectivel…
cs.CV2023★ 5 cited
FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection
Chunyong Hu, Hang Zheng, Kun Li +11
Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features in…