20 citations · 23 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
ViTAR: Vision Transformer with Any Resolution
Qihang Fan, Quanzeng You, Xiaotian Han +5
This paper tackles a significant challenge faced by Vision Transformers (ViTs): their constrained scalability across different image resolutions. Typically, ViTs experience a perfo…
cs.CL2024★ 20 cited
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Yiqi Wang, Wentao Chen, Xiaotian Han +7
Strong Artificial Intelligence (Strong AI) or Artificial General Intelligence (AGI) with abstract reasoning ability is the goal of next-generation AI. Recent advancements in Large…
cs.CV2024
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
Xiaotian Han, Yiqi Wang, Bohan Zhai +2
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. Visual instruction fine-tuning (IFT) is a vital process for aligning M…