activity
20242026
most citedMM-RLHF: The Next Step Forward in Multimodal LLM Alignment

1 citations · 1 across the 22 of their papers we have counts for

collaborators
Showing cs.CVShow all

18 papers · 1 filter

cs.CV2026

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

Jiuzhou Lin, Junlong Wu, Fei Zuo +11

Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typical…

cs.CV2026

Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

Junlong Wu, Jiuzhou Lin, Jia Sun +7

Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text…

cs.CV2026

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

Honglie Wang, Jia Sun, Zijun Li +11

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recen…

cs.CV2026

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation

Zijun Li, Yimin Zhou, Jia Sun +12

Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when…

cs.CV2026

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

Yankai Yang, Yancheng Long, Bin Wen +4

Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos sha…

cs.CV2026

SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

Yankai Yang, Yancheng Long, Wei Chen +7

Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which…