19 papers
When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution
Yu Shi, Yuyao Zhang, Yu-wing Tai
Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In…
IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning
Yuening Cai, Junwei Zhou, Youran Qu +1
Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary axis of spatial competence u…
Beyond MoCap: Scaling Motion Tokenizers with Synthetic Human Motion for Generative Modeling
Yiwen Yan, Wanning He, Yu-Wing Tai
Human motion generation models are fundamentally constrained by the limited diversity of motion capture datasets, which predominantly contain common, repetitive actions and fail to…
HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing
Yuyao Zhang, Alexander Huang-Menders, Yu-Wing Tai
High-resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion-based editors remain computationally inefficient and constr…
AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling
Ziyang Mai, Yuyao Zhang, Yu-Wing Tai
Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (UHR) long videos remains pro…
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang
Reasoning Video Object Segmentation is a challenging task, aiming at generating a mask sequence from an input video given a complex and implicit text query. While existing works fi…