papers

Publications (19)

cs.CV2026

Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning

Xinyu Liu, Hangjie Yuan, Yujie Wei +9

Unified video models exhibit strong capabilities in understanding and generation, yet they struggle with reason-informed visual editing even when equipped with powerful internal vi…

cs.CV2025

Exploiting Discriminative Codebook Prior for Autoregressive Image Generation

Longxiang Tang, Ruihang Chu, Xiang Wang +6

Advanced discrete token-based autoregressive image generation systems first tokenize images into sequences of token indices with a codebook, and then model these sequences in an au…

cs.CV2025

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?

Yujin Han, Andi Han, Wei Huang +2

Despite the remarkable success of diffusion models (DMs) in data generation, they exhibit specific failure cases with unsatisfactory outputs. We focus on one such limitation: the a…

cs.LG2024

Conformalized Semi-supervised Random Forest for Classification and Abnormality Detection

Yujin Han, Mingwenchan Xu, Leying Guan

The Random Forests classifier, a widely utilized off-the-shelf classification tool, assumes training and test samples come from the same distribution as other standard classifiers.…

cs.CV2026

Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation

Chika Maduabuchi, Hao Chen, Yujin Han +1

Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where…

cs.CV2026

Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance

Yujie Wei, Shiwei Zhang, Hangjie Yuan +8

Mixture-of-Experts (MoE) has emerged as a powerful paradigm for scaling model capacity while preserving computational efficiency. Despite its notable success in large language mode…

cs.LG2025

Capturing Conditional Dependence via Auto-regressive Diffusion Models

Xunpeng Huang, Yujin Han, Difan Zou +2

Diffusion models have demonstrated appealing performance in both image and video generation. However, many works discover that they struggle to capture important, high-level relati…

cs.LG2024

Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference

Yujin Han, Difan Zou

Standard empirical risk minimization (ERM) models may prioritize learning spurious correlations between spurious features and true labels, leading to poor accuracy on groups where…

cs.LG2023

GAN-based Vertical Federated Learning for Label Protection in Binary Classification

Yujin Han, Leying Guan

Split learning (splitNN) has emerged as a popular strategy for addressing the high computational costs and low modeling efficiency in Vertical Federated Learning (VFL). However, de…

cs.CV2026

AesRM: Improving Video Aesthetics with Expert-Level Feedback

Yujin Han, Yujie Wei, Yefei He +7

Despite rapid advances in photorealistic video generation, real-world applications such as filmmaking require video aesthetics, e.g., harmonious colors and cinematic lighting, beyo…

cs.CV2025

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

Hao Chen, Yujin Han, Fangyi Chen +7

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for…

cs.CV2025

Parallelized Autoregressive Visual Generation

Yuqing Wang, Shuhuai Ren, Zhijie Lin +6

Autoregressive models have emerged as a powerful approach for visual generation but suffer from slow inference speed due to their sequential token-by-token prediction process. In t…

cs.CV2026

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Yujie Wei, Yujin Han, Zhekai Chen +20

Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier mo…

cs.LG2022

FedGBF: An efficient vertical federated learning framework via gradient boosting and bagging

Yujin Han, Pan Du, Kai Yang

Federated learning, conducive to solving data privacy and security problems, has attracted increasing attention recently. However, the existing federated boosting model sequentiall…

cs.CL2024

Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability

Yujin Han, Lei Xu, Sirui Chen +2

Large language models (LLMs) have shown remarkable capability in natural language tasks, yet debate persists on whether they truly comprehend deep structure (i.e., core semantics)…

cs.CV2026

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Xinyu Liu, Shihao Li, Weihong Lin +10

The paper introduces ReBind, a framework that uses structured instructions with explicit reference tokens to improve multi‑reference image‑conditioned video editing, enabling preci…

#multi-reference video editing#diffusion models#structured instructions#reference tokens
cs.AI2026

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

Zichao Yu, Chengzhi Yu, Shengze Xu +4

On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repet…

cs.CV2024

Slight Corruption in Pre-training Data Makes Better Diffusion Models

Hao Chen, Yujin Han, Diganta Misra +6

Diffusion models (DMs) have shown remarkable capabilities in generating realistic high-quality images, audios, and videos. They benefit significantly from extensive pre-training on…

cs.CL2025

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han +5

Although unified MLLMs aim to unify generation and understanding, they are considered to exhibit an internal gap, with understanding outperforming generation. Through large-scale e…