From the 1 of 4 linked papers with an AI index.
4 papers
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery
Xinyue Xu, Zheng Zhang, Kunyang Ma +5
The paper introduces DM-KG, a direction‑metric knowledge graph that extracts 3D spatial relationships from street‑view images and injects them into vision‑language models to improv…
EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents
Sai Ma, Zhuang Li, Sichao Li +4
Earth Observation (EO) analysis is inherently interactive: resolving uncertainty often requires expanding the region of interest, retrieving historical observations, and switching…
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
Junxian Li, Xinyue Xu, Sai Ma +2
Multimodal Large Language Models (MLLMs) frequently suffer from unfaithfulness, generating reasoning chains that drift from visual evidence or contradict final predictions. We prop…
Evaluating LLM Understanding via Structured Tabular Decision Simulations
Sichao Li, Xinyue Xu, Xiaomeng Li
Large language models (LLMs) often achieve impressive predictive accuracy, yet correctness alone does not imply genuine understanding. True LLM understanding, analogous to human ex…