8 papers · 1 filter
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models
Jiazi Bu, Pengyang Ling, Yujie Zhou +10
Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given…
The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
Yuxi Wang, Chengkai Jin, Yufei Liu +6
4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector that fails under heavy occl…
AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO
Jiazi Bu, Pengyang Ling, Yujie Zhou +8
Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that t…
StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation
Yinxi He, Kang Liao, Chunyu Lin +2
This paper introduces StructDiff, a generative framework based on a single-scale diffusion model for single-image generation. Single-image generation aims to synthesize diverse sam…
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
Jiazi Bu, Pengyang Ling, Yujie Zhou +8
Group Relative Policy Optimization (GRPO) has emerged as a powerful framework for preference alignment in text-to-image (T2I) flow models. However, we observe that the standard par…
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
Yuxi Wang, Wenqi Ouyang, Tianyi Wei +3
Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency,…