3 papers
cs.CV2026
When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution
Yu Shi, Yuyao Zhang, Yu-wing Tai
Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In…
cs.CV2025
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
Alexander Huang-Menders, Xinhang Liu, Andy Xu +3
SmartAvatar is a vision-language-agent-driven framework for generating fully rigged, animation-ready 3D human avatars from a single photo or textual prompt. While diffusion-based m…
cs.LG2025
LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration
Yuyao Zhang, Jinghao Li, Yu-Wing Tai
Text-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing.…