5 papers · 1 filter
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
Xuelu Feng, Yunsheng Li, Ziyu Wan +4
Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key challenge, however, lies in desi…
Olympus: A Universal Task Router for Computer Vision Tasks
Yuanze Lin, Yunsheng Li, Dongdong Chen +3
We introduce Olympus, a new approach that transforms Multimodal Large Language Models (MLLMs) into a unified framework capable of handling a wide array of computer vision tasks. Ut…
Benchmarking Large and Small MLLMs
Xuelu Feng, Yunsheng Li, Dongdong Chen +4
Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior qua…
Pluralistic Salient Object Detection
Xuelu Feng, Yunsheng Li, Dongdong Chen +4
We introduce pluralistic salient object detection (PSOD), a novel task aimed at generating multiple plausible salient segmentation results for a given input image. Unlike conventio…
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
Zixin Zhu, Xuelu Feng, Dongdong Chen +3
In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent r…