6 papers
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
Hongji Yang, Songlian Li, Yucheng Zhou +4
Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in…
HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
Hongji Yang, Yucheng Zhou, Wencheng Han +3
While Large Vision-Language Models (VLMs) demonstrate impressive general visual capabilities, they remain artistically blind and unable to offer professional evaluation of artworks…
HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
Hongji Yang, Yucheng Zhou, Wencheng Han +4
Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts invo…
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
Hongji Yang, Yucheng Zhou, Wencheng Han +1
Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this,…
DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models
Hongji Yang, Wencheng Han, Yucheng Zhou +1
In this paper, we introduce DC (Decouple)-ControlNet, a highly flexible and precisely controllable framework for multi-condition image generation. The core idea behind DC-ControlNe…
Retrieval-guided Cross-view Image Synthesis
Hongji Yang, Yiru Li, Yingying Zhu
Information retrieval techniques have demonstrated exceptional capabilities in identifying semantic similarities across diverse domains through robust feature representations. Howe…