1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
Shaokai Ye, Vasileios Saveris, Yihao Qian +3
Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large langu…
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang +7
Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing ben…
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Roman Bachmann, Jesse Allardice, David Mizrahi +6
Image tokenization has enabled major advances in autoregressive image generation by providing compressed, discrete representations that are more efficient to process than raw pixel…
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Elmira Amirloo, Jean-Philippe Fauconnier, Christoph Roesmann +8
Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains…