collaborators

15 papers

cs.CY2026

From Causal Discovery to Implementation: An Agentic AI Framework for E-Scooter Mobility Hub Planning Across 29 German Cities

Meng Jin, Melanie Handrich, Simone Martinenz +2

Existing approaches to e-scooter mobility hub planning lack city-type-specific causal evidence. Demand models are typically correlational, built on proprietary trip data, and do no…

cs.CV2026

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

Keqin Zeng, Shuting Su, Shihao Lin +2

Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data…

cs.LG2026

Multi-Modal Contrastive Learning for Implicit Earth Embeddings via Location Tying

Jonathan Hecht, Lukas Arzoumanidis, Ziyue Li +1

Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations. To overcome this challenge, self-supervised pre-training is a possible solu…

cs.CL2026

SQLBench: A Comprehensive Evaluation for Text-to-SQL Capabilities of Large Language Models

Bin Zhang, Yuxiao Ye, Guoqing Du +8

Large Language Models (LLMs) have emerged as a powerful tool in advancing the Text-to-SQL task, significantly outperforming traditional methods.Nevertheless, as a nascent research…

cs.CV2026

GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models

Qinghongbing Xie, Zhaoyuan Xia, Feng Zhu +4

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for autonomous driving, embodied AI and general AI. Exist…

cs.CV2025

Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model

Jihua Peng, Qianxiong Xu, Yichen Liu +4

Group activity detection (GAD) aims to simultaneously identify group members and categorize their collective activities within video sequences. Existing deep learning-based methods…