6 papers
OmniPrism: Learning Disentangled Visual Concept for Image Generation
Yangyang Li, Daqing Liu, Wu Liu +4
Creative visual concept generation often draws inspiration from specific concepts in a reference image to produce relevant outcomes. However, existing methods are typically constra…
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
Chen Chen, Jiawei Shao, Dakuan Lu +4
Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot vis…
ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion
Wei Huang, Peining Li, Meiyu Liang +7
Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. H…
CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
Disheng Liu, Yiran Qiao, Wuche Liu +5
True intelligence hinges on the ability to uncover and leverage hidden causal relations. Despite significant progress in AI and computer vision (CV), there remains a lack of benchm…
HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
Kun Liu, Qi Liu, Xinchen Liu +5
Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely gener…
LMAgent: A Large-scale Multimodal Agents Society for Multi-user Simulation
Yijun Liu, Wu Liu, Xiaoyan Gu +3
The believable simulation of multi-user behavior is crucial for understanding complex social systems. Recently, large language models (LLMs)-based AI agents have made significant p…