4 citations · 4 across the 5 of their papers we have counts for
3 papers · 1 filter
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
Shaokai Ye, Vasileios Saveris, Yihao Qian +3
Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large langu…
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Rui Tian, Mingfei Gao, Mingze Xu +5
We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centr…
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi +6
Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform…