activity
20242026
most cited4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

4 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2026

Distributionally Robust Token Optimization in RLHF

Yeping Jin, Jiaming Hu, Ioannis Ch. Paschalidis

Large Language Models (LLMs) tend to respond correctly to prompts that align well with the data they were trained and fine-tuned on. Yet, small shifts in wording, format, or langua…

cs.CV2026

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning

Shaokai Ye, Vasileios Saveris, Yihao Qian +3

Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large langu…

cs.LG2026

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium

Jiaming Hu, Jiamu Bai, Haoyu Wang +2

Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a mainstream branch of RLHF, Direct…

cs.CV2025

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Rui Tian, Mingfei Gao, Mingze Xu +5

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centr…

cs.CV20244 cited

4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi +6

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform…