activity
20212025
most citedA Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks

19 citations · 54 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Adaptive Critical Subgraph Mining for Cognitive Impairment Conversion Prediction with T1-MRI-based Brain Network

Yilin Leng, Wenju Cui, Bai Chen +3

Prediction the conversion to early-stage dementia is critical for mitigating its progression but remains challenging due to subtle cognitive impairments and structural brain change…

cs.CV2023

Review of Large Vision Models and Visual Prompt Engineering

Jiaqi Wang, Zhengliang Liu, Lin Zhao +18

Visual prompt engineering is a fundamental technology in the field of visual and image Artificial General Intelligence, serving as a key component for achieving zero-shot capabilit…

cs.CV20237 cited

Instruction-ViT: Multi-Modal Prompts for Instruction Learning in ViT

Zhenxiang Xiao, Yuzhong Chen, Lu Zhang +14

Prompts have been proven to play a crucial role in large language models, and in recent years, vision models have also been using prompts to improve scalability for multiple downst…

cs.CV20221 cited

Eye-gaze-guided Vision Transformer for Rectifying Shortcut Learning

Chong Ma, Lin Zhao, Yuzhong Chen +15

Learning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning the meaningful and useful representations, thus jeopardizing the gen…

cs.CV202210 cited

Mask-guided Vision Transformer (MG-ViT) for Few-Shot Learning

Yuzhong Chen, Zhenxiang Xiao, Lin Zhao +10

Learning with little data is challenging but often inevitable in various application scenarios where the labeled data is limited and costly. Recently, few-shot learning (FSL) gaine…