6 papers
Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
Qing Zhao, Weijian Deng, Pengxu Wei +1
Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over time remains challenging. Dynam…
Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection
Zijie Cao, Weijie Tu, Yao Xiao +3
Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, t…
Thinking Before You Speak: A Proactive Test-time Scaling Approach
Cong Liu, Wenchang Chai, Hejun Wu +3
Large Language Models (LLMs) often exhibit deficiencies with complex reasoning tasks, such as maths, which we attribute to the discrepancy between human reasoning patterns and thos…
Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention into Convolutions
ZiYi Dong, Chengxing Zhou, Weijian Deng +3
Contemporary diffusion models built upon U-Net or Diffusion Transformer (DiT) architectures have revolutionized image generation through transformer-based attention mechanisms. The…
Decoder-Only LLMs are Better Controllers for Diffusion Models
Ziyi Dong, Yao Xiao, Pengxu Wei +1
Groundbreaking advancements in text-to-image generation have recently been achieved with the emergence of diffusion models. These models exhibit a remarkable ability to generate hi…
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
Ziyi Dong, Pengxu Wei, Liang Lin
State-of-the-arts text-to-image generation models such as Imagen and Stable Diffusion Model have succeed remarkable progresses in synthesizing high-quality, feature-rich images wit…