collaborators

6 papers

cs.CV2026

When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution

Yu Shi, Yuyao Zhang, Yu-wing Tai

Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In…

cs.CV2026

UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention

Yuyao Zhang, Yu-Wing Tai

Ultra-high-resolution text-to-image generation is increasingly vital for applications requiring fine-grained textures and global structural fidelity, yet state-of-the-art text-to-i…

cs.CV2026

HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing

Yuyao Zhang, Alexander Huang-Menders, Yu-Wing Tai

High-resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion-based editors remain computationally inefficient and constr…

cs.CV2026

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling

Ziyang Mai, Yuyao Zhang, Yu-Wing Tai

Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (UHR) long videos remains pro…

cs.LG2025

LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration

Yuyao Zhang, Jinghao Li, Yu-Wing Tai

Text-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing.…

cs.CV2025

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents

Alexander Huang-Menders, Xinhang Liu, Andy Xu +3

SmartAvatar is a vision-language-agent-driven framework for generating fully rigged, animation-ready 3D human avatars from a single photo or textual prompt. While diffusion-based m…