2 papers
cs.CV2026
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
Shaokai Ye, Vasileios Saveris, Yihao Qian +3
Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large langu…
cs.CV2026
LLaVAction: evaluating and training multi-modal large language models for action understanding
Haozhe Qi, Shaokai Ye, Alexander Mathis +1
Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multim…