31 citations · 197 across the 130 of their papers we have counts for
6 papers · 1 filter
OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation
Yajing Xu, Yarong Lan, Jiaoyan Chen +6
While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rely on coarse-grained descriptio…
Generic Interpretation Approach for Transformer Models Incorporating Heterogenous Attention Structures
Yongjin Cui, Xiaohui Fan, Huajun Chen
Transformer has significantly propelled the development of artificial intelligence, and certainly the development of agents as well. We categorize attention structures of Transform…
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
Yichi Zhang, Zhuo Chen, Lingbing Guo +2
Understanding and reasoning with abstractive information from the visual modality presents significant challenges for current multi-modal large language models (MLLMs). Among the v…
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
Yichi Zhang, Zhuo Chen, Lingbing Guo +4
Multi-modal large language models (MLLMs) incorporate heterogeneous modalities into LLMs, enabling a comprehensive understanding of diverse scenarios and objects. Despite the proli…
InstructRL4Pix: Training Diffusion for Image Editing by Reinforcement Learning
Tiancheng Li, Jinxiu Liu, Huajun Chen +1
Instruction-based image editing has made a great process in using natural human language to manipulate the visual content of images. However, existing models are limited by the qua…
Knowledge Perceived Multi-modal Pretraining in E-commerce
Yushan Zhu, Huaixiao Tou, Wen Zhang +4
In this paper, we address multi-modal pretraining of product data in the field of E-commerce. Current multi-modal pretraining methods proposed for image and text modalities lack ro…