4 papers · 1 filter
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
Haoyuan Sun, Jing Wang, Yuxin Song +9
Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for furth…
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
Xinyao Zhang, Wenkai Dong, Yuxin Song +10
Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely…
CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
YuXin Song, Yu Lu, Haoyuan Sun +6
Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for…
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
Ziliang Gan, Yu Lu, Dong Zhang +9
In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiariti…