2 papers
cs.CV2025
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Boqiang Zhang, Kehan Li, Zesen Cheng +12
In this paper, we propose VideoLLaMA3, a more advanced multimodal foundation model for image and video understanding. The core design philosophy of VideoLLaMA3 is vision-centric. T…
eess.IV2024
Large-vocabulary forensic pathological analyses via prototypical cross-modal contrastive learning
Chen Shen, Chunfeng Lian, Wanqing Zhang +11
Forensic pathology is critical in determining the cause and manner of death through post-mortem examinations, both macroscopic and microscopic. The field, however, grapples with is…