activity
20232025
collaborators

5 papers

cs.CV2025

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

Penglei Sun, Yaoxian Song, Xiangru Zhu +7

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…

cs.CL2024

Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective

Xiangru Zhu, Penglei Sun, Yaoxian Song +6

Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from…

cs.IR2024

Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval

Haoyu Liu, Yaoxian Song, Xuwu Wang +4

With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is need…

cs.CV2023

A Contrastive Compositional Benchmark for Text-to-Image Synthesis: A Study with Unified Text-to-Image Fidelity Metrics

Xiangru Zhu, Penglei Sun, Chengyu Wang +4

Text-to-image (T2I) synthesis has recently achieved significant advancements. However, challenges remain in the model's compositionality, which is the ability to create new combina…

cs.AI2023

M^2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge Base

Zhiwei Zha, Jiaan Wang, Zhixu Li +3

Multimodal knowledge bases (MMKBs) provide cross-modal aligned knowledge crucial for multimodal tasks. However, the images in existing MMKBs are generally collected for entities in…