30 papers
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Yicheng Xiao, Wenxun Dai, Xinran Qin +22
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present…
Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs
Mingyu Wang, Weilin Jin, Wenbo Li +5
Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is inconsistent with or unsupported…
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
Yong Liu, Xiaolong Fu, Zihang Xu +10
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for t…
Self Gradient Forcing: Native Long Video Extrapolation
Junhao Zhuang, Shiyi Zhang, Yuxuan Bian +11
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-tru…
JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid
Peidong Liu, Yongce Liu, Songyan Guo +34
JoyAI-Sim is a toolchain that connects real robots, simulation, and human demonstrations to enable scalable evaluation and generation of robot training data using calibrated digita…
Perceptual Flow Matching for Few-Step Generative Modeling
Chuyang Zhao, Yifei Song, Hongfa Wang +7
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conve…