2 papers
cs.CV2026
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
Kaiwen Xue, Tao Wei, Guoxin Zhang +5
Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluat…
cs.AI2025
CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
Kaiwen Xue, Chenglong Li, Zhonghong Ou +10
Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. T…