18 papers
InstanceControl: Controllable Complex Image Generation without Instance Labeling
Xiaoyu Liu, Huan Wang, Fan Li +4
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. Howev…
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Chonghuinan Wang, Zhikai Chen, Chunwei Wang +9
The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving…
ShotCrop: Cropping Human-Centric Images into Cinematic Triple-Shot Compositions
Dehong Kong, Lina Lei, Lingtao Zheng +10
Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots from one scene. In practice…
Fast Image Super-Resolution via Consistency Rectified Flow
Jiaqi Xu, Wenbo Li, Haoze Sun +8
Diffusion models (DMs) have demonstrated remarkable success in real-world image super-resolution (SR), yet their reliance on time-consuming multi-step sampling largely hinders thei…
PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models
Yongsen Cheng, Kai Liu, Kaiwen Tao +5
Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained sc…
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
Zhixin Wang, Jiaming Xu, Tianyi Zhou +10
Effectively scaling Reinforcement Learning (RL) is crucial for enhancing the reasoning and alignment of Large Language Models. The massive data and complex execution flows inherent…