3 papers
cs.CV2026
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
Yuhe Liu, Zhenxiong Tan, Yujia Hu +2
Recent advances in diffusion-based controllable visual generation have led to remarkable improvements in image quality. However, these powerful models are typically deployed on clo…
cs.CV2025
Image Editing As Programs with Diffusion Models
Yujia Hu, Songhua Liu, Zhenxiong Tan +2
While diffusion models have achieved remarkable success in text-to-image generation, they encounter significant challenges with instruction-driven image editing. Our research highl…
cs.CV2025
Flash Sculptor: Modular 3D Worlds from Objects
Yujia Hu, Songhua Liu, Xingyi Yang +1
Existing text-to-3D and image-to-3D models often struggle with complex scenes involving multiple objects and intricate interactions. Although some recent attempts have explored suc…