activity
20242026
collaborators

5 papers

cs.CV2026

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo +15

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only sing…

cs.LG2025

Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment

Yawen Shao, Jie Xiao, Kai Zhu +4

Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO…

cs.CV2025

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan +13

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…

cs.CV2025

MangaNinja: Line Art Colorization with Precise Reference Following

Zhiheng Liu, Ka Leong Cheng, Xi Chen +7

Derived from diffusion models, MangaNinjia specializes in the task of reference-guided line art colorization. We incorporate two thoughtful designs to ensure precise character deta…

cs.AI2024

The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

Ruili Feng, Han Zhang, Zhantao Yang +7

We present The Matrix, the first foundational realistic world simulator capable of generating continuous 720p high-fidelity real-scene video streams with real-time, responsive cont…