Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
Chen Chen, Jiawei Shao, Dakuan Lu +4
Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot vis…
cs.AI2026
ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion
Wei Huang, Peining Li, Meiyu Liang +7
Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. H…
cs.AI2024
LMAgent: A Large-scale Multimodal Agents Society for Multi-user Simulation
Yijun Liu, Wu Liu, Xiaoyan Gu +3
The believable simulation of multi-user behavior is crucial for understanding complex social systems. Recently, large language models (LLMs)-based AI agents have made significant p…