12 citations · 17 across the 2 of their papers we have counts for
2 papers
cs.CV2025★ 5 cited
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Boqiang Zhang, Kehan Li, Zesen Cheng +12
In this paper, we propose VideoLLaMA3, a more advanced multimodal foundation model for image and video understanding. The core design philosophy of VideoLLaMA3 is vision-centric. T…
eess.IV2024★ 12 cited
Large-vocabulary forensic pathological analyses via prototypical cross-modal contrastive learning
Chen Shen, Chunfeng Lian, Wanqing Zhang +11
Forensic pathology is critical in determining the cause and manner of death through post-mortem examinations, both macroscopic and microscopic. The field, however, grapples with is…