activity
20242026
collaborators

5 papers

cs.CV2026

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

Xinyue Xu, Zheng Zhang, Kunyang Ma +5

As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street v…

cs.RO2026

Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge

Li Kang, Heng Zhou, Xiufeng Song +41

Recent advancements in multimodal large language models and vision-languageaction models have significantly driven progress in Embodied AI. As the field transitions toward more com…

cs.DC2026

Self-Evolving Distributed Memory Architecture for Scalable AI Systems

Zixuan Li, Chuanzhen Wang, Haotian Sun

Distributed AI systems face critical memory management challenges across computation, communication, and deployment layers. RRAM based in memory computing suffers from scalability…

cs.AI2025

BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models

Thierry Blankenstein, Jialin Yu, Zixuan Li +6

Agents backed by large language models (LLMs) increasingly rely on external tools drawn from marketplaces where multiple providers offer functionally equivalent options. This raise…

cs.CV2024

GPT-4V Explorations: Mining Autonomous Driving

Zixuan Li

This paper explores the application of the GPT-4V(ision) large visual language model to autonomous driving in mining environments, where traditional systems often falter in underst…