3 papers
cs.CV2025
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
Penglei Sun, Yaoxian Song, Xiangru Zhu +7
Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…
cs.CL2025
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
Xiangru Zhu, Penglei Sun, Yaoxian Song +6
Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from…
cs.AI2025
M^2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge Base
Zhiwei Zha, Jiaan Wang, Zhixu Li +3
Multimodal knowledge bases (MMKBs) provide cross-modal aligned knowledge crucial for multimodal tasks. However, the images in existing MMKBs are generally collected for entities in…