34 citations · 76 across the 25 of their papers we have counts for
24 papers · 1 filter
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
Dengyang Jiang, Ruoyi Du, Zhennan Chen +10
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on smal…
Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
Xin Jin, Huanqia Cai, Zhen Li +9
Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric scores than as a deterministic sc…
coDrawAgents: A Multi-Agent Dialogue Framework for Compositional Image Generation
Chunhan Li, Qifeng Wu, Jia-Hui Pan +7
Text-to-image generation has advanced rapidly, but existing models still struggle with faithfully composing multiple objects and preserving their attributes in complex scenes. We p…
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
Sicong Leng, Jing Wang, Jiaxi Li +12
Large multimodal reasoning models have achieved rapid progress, but their advancement is constrained by two major limitations: the absence of open, large-scale, high-quality long c…
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
Yuming Jiang, Siteng Huang, Shengke Xue +10
This paper presents RynnVLA-001, a vision-language-action(VLA) model built upon large-scale video generative pretraining from human demonstrations. We propose a novel two-stage pre…
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
Weichen Fan, Chenyang Si, Junhao Song +16
We present Vchitect-2.0, a parallel transformer architecture designed to scale up video diffusion models for large-scale text-to-video generation. The overall Vchitect-2.0 system h…