3 papers
cs.CV2026
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Jinbo Yan, Limeng Qiao, Jie Qin +3
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine…
cs.CV2024
DiffusionAgent: Navigating Expert Models for Agentic Image Generation
Jie Qin, Jie Wu, Weifeng Chen +1
In the accelerating era of human-instructed visual content creation, diffusion models have demonstrated remarkable generative potential. Yet their deployment is constrained by a du…
cs.CV2023
DiffusionEngine: Diffusion Model is Scalable Data Engine for Object Detection
Manlin Zhang, Jie Wu, Yuxi Ren +7
Data is the cornerstone of deep learning. This paper reveals that the recently developed Diffusion Model is a scalable data engine for object detection. Existing methods for scalin…