20 citations · 23 across the 3 of their papers we have counts for
7 papers
Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
Haogeng Liu, Quanzeng You, Xiaotian Han +4
In the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs).…
ViTAR: Vision Transformer with Any Resolution
Qihang Fan, Quanzeng You, Xiaotian Han +5
This paper tackles a significant challenge faced by Vision Transformers (ViTs): their constrained scalability across different image resolutions. Typically, ViTs experience a perfo…
InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding
Haogeng Liu, Quanzeng You, Xiaotian Han +7
Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intri…
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Yiqi Wang, Wentao Chen, Xiaotian Han +7
Strong Artificial Intelligence (Strong AI) or Artificial General Intelligence (AGI) with abstract reasoning ability is the goal of next-generation AI. Recent advancements in Large…
InfiMM-Eval: Complex Open-Ended Reasoning Evaluation For Multi-Modal Large Language Models
Xiaotian Han, Quanzeng You, Yongfei Liu +9
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. These models not only excel in traditional vision-language tasks but a…
Improving In-Context Learning in Diffusion Models with Visual Context-Modulated Prompts
Tianqi Chen, Yongfei Liu, Zhendong Wang +4
In light of the remarkable success of in-context learning in large language models, its potential extension to the vision domain, particularly with visual foundation models like St…