collaborators

6 papers

cs.CV2025

From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition

Ling Lo, Kelvin C. K. Chan, Wen-Huang Cheng +1

Existing models often struggle with complex temporal changes, particularly when generating videos with gradual attribute transitions. The most common prompt interpolation approach…

cs.CV2025

HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis

Xiaoyuan Wang, Yizhou Zhao, Botao Ye +6

We propose HoliGS, a novel deformable Gaussian splatting framework that addresses embodied view synthesis from long monocular RGB videos. Unlike prior 4D Gaussian splatting and dyn…

cs.CV2025

CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts

Lee Hsin-Ying, Kelvin C. K. Chan, Ming-Hsuan Yang

While text-to-image generative models can synthesize diverse and faithful content, subject variation across multiple generations limits their application to long-form content gener…

cs.CV2024

HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior

Li-Yuan Tsao, Hao-Wei Chen, Hao-Wei Chung +4

Text-to-image diffusion models have emerged as powerful priors for real-world image super-resolution (Real-ISR). However, existing methods may produce unintended results due to noi…

cs.CV2024

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton +8

Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it rem…

cs.CV2024

A Simple Approach to Unifying Diffusion-based Conditional Generation

Xirui Li, Charles Herrmann, Kelvin C. K. Chan +4

Recent progress in image generation has sparked research into controlling these models through condition signals, with various methods addressing specific challenges in conditional…