activity
20232026
collaborators

5 papers

cs.CV2026

Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation

Ernie Chu, Vishal M. Patel

Diffusion Transformers (DiTs) have achieved state-of-the-art video generation quality, but they incur immense computational cost because standard inference applies the same number…

cs.CV2026

Face-to-Face: A Video Dataset for Multi-Person Interaction Modeling

Ernie Chu, Vishal M. Patel

Modeling the reactive tempo of human conversation remains difficult because most audio-visual datasets portray isolated speakers delivering short monologues. We introduce \textbf{F…

eess.IV2024

FIPER: Factorized Features for Robust Image Super-Resolution and Compression

Yang-Che Sun, Cheng Yu Yeo, Ernie Chu +2

In this work, we propose using a unified representation, termed Factorized Features, for low-level vision tasks, where we test on Single Image Super-Resolution (SISR) and \textbf{I…

cs.CV2024

Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion Models

Chun-Yen Shih, Li-Xuan Peng, Jia-Wei Liao +3

Diffusion Models have emerged as powerful generative models for high-quality image synthesis, with many subsequent image editing techniques based on them. However, the ease of text…

eess.AS2023

Deep Complex U-Net with Conformer for Audio-Visual Speech Enhancement

Shafique Ahmed, Chia-Wei Chen, Wenze Ren +7

Recent studies have increasingly acknowledged the advantages of incorporating visual data into speech enhancement (SE) systems. In this paper, we introduce a novel audio-visual SE…