9 papers
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
Yunze Tong, Mushui Liu, Canyu Zhao +9
With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align t…
RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
Yuhan Li, Fangao Zeng, Sicong Kang +5
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequenti…
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
Yunze Tong, Mushui Liu, Canyu Zhao +7
Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome-based reward to all preceding d…
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Peng Zhang, Guanghao Zhang, Wanggui He +10
Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
Longxiang Zhang, Weilong Dai, Guanghao Zhang +2
Multimodal large language models (MLLMs) have emerged as a powerful backbone for multimodal embeddings. Recent methods introduce chain-of-thought (CoT) reasoning into the embedding…
simpleposter: A simple baseline for product poster generation
Benlei Cui, Fangao Zeng, Weitao Jiang +6
Product poster generation poses distinct challenges beyond general poster design, requiring both faithful preservation of product appearance and precise control over dense, multi-l…