activity
20232026
most citedMultimodal Foundation Models: From Specialists to General-Purpose Assistants

25 citations · 57 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CV2026

UniT: Unified Multimodal Chain-of-Thought Test-time Scaling

Leon Liangyu Chen, Haoyu Ma, Zhipeng Fan +11

Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single pass without iteratively refining their…

cs.CV2024

Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment

Xin Xiao, Bohong Wu, Jiacong Wang +3

Existing image-text modality alignment in Vision Language Models (VLMs) treats each text token equally in an autoregressive manner. Despite being simple and effective, this method…

cs.CV202314 cited

Large Language Models are Visual Reasoning Coordinators

Liangyu Chen, Bo Li, Sheng Shen +5

Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsen…

cs.CV2023

Improved Baselines with Visual Instruction Tuning

Haotian Liu, Chunyuan Li, Yuheng Li +1

Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal co…

cs.CV2023

HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Bohan Zhai, Shijia Yang, Chenfeng Xu +4

Current Large Multimodal Models (LMMs) achieve remarkable progress, yet there remains significant uncertainty regarding their ability to accurately apprehend visual details, that i…

cs.CV202312 cited

Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao +9

Large Multimodal Models (LMM) are built across modalities and the misalignment between two modalities can result in "hallucination", generating textual outputs that are not grounde…