4 papers
Unifying Distributional Training for One-Step Visual Generation
Chi Zhang, Shi Haoyang, Haoyang Shi +10
Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unifi…
From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
Chi Zhang, Yueyi Liu, Shi Haoyang +6
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We…
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
Chi Zhang, Haoyang Shi, Yueyi Liu +4
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatar…
Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning
Ziyi Chen, Haoyan Shi, Sunhan Xu +1
Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pairs. Prior works either pred…