2 citations · 4 across the 6 of their papers we have counts for
6 papers
StableI2I: Spotting Unintended Changes in Image-to-Image Transition
Jiayang Li, Shuo Cao, Xiaohui Li +6
In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. H…
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
Zhipeng Huang, Shaobin Zhuang, Canmiao Fu +7
Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack th…
A General Theory for Compositional Generalization
Jingwen Fu, Zhizheng Zhang, Yan Lu +1
Compositional Generalization (CG) embodies the ability to comprehend novel combinations of familiar concepts, representing a significant cognitive leap in human intellectual advanc…
VisualCritic: Making LMMs Perceive Visual Quality Like Humans
Zhipeng Huang, Zhizheng Zhang, Yiting Lu +3
At present, large multimodal models (LMMs) have exhibited impressive generalization capabilities in understanding and generating visual signals. However, they currently still lack…
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
Zhipeng Huang, Zhizheng Zhang, Zheng-Jun Zha +2
The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very…
Reinforced UI Instruction Grounding: Towards a Generic UI Task Automation API
Zhizheng Zhang, Wenxuan Xie, Xiaoyi Zhang +1
Recent popularity of Large Language Models (LLMs) has opened countless possibilities in automating numerous AI tasks by connecting LLMs to various domain-specific models or APIs, w…