55 citations · 55 across the 1 of their papers we have counts for
5 papers · 1 filter
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
Yuxuan Cai, Jiangning Zhang, Zhenye Gan +9
Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehen…
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
Yuxuan Cai, Jiangning Zhang, Haoyang He +7
The success of Large Language Models (LLMs) has inspired the development of Multimodal Large Language Models (MLLMs) for unified understanding of vision and language. However, the…
LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation
Guoping Xu, Xingrong Wu, Xuan Zhang +1
Medical image segmentation plays an essential role in developing computer-assisted diagnosis and therapy systems, yet still faces many challenges. In the past few years, the popula…
View N-gram Network for 3D Object Retrieval
Xinwei He, Tengteng Huang, Song Bai +1
How to aggregate multi-view representations of a 3D object into an informative and discriminative one remains a key challenge for multi-view 3D object retrieval. Existing methods e…
Triplet-Center Loss for Multi-View 3D Object Retrieval
Xinwei He, Yang Zhou, Zhichao Zhou +2
Most existing 3D object recognition algorithms focus on leveraging the strong discriminative power of deep learning models with softmax loss for the classification of 3D data, whil…