From the 1 of 11 linked papers with an AI index.
11 papers
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
The paper investigates whether large language models (LLMs) have distinct planning abilities by applying multidimensional item response theory to benchmark data, uncovering two sep…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fucai Ke, Joy Hsu, Zhixi Cai +10
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…
ARIS: Agentic and Relationship Intelligence System for Social Robots
Stavya Datta, Fucai Ke, Leimin Tian +1
Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still struggle with multi-turn enga…
Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
Instruction granularity is an important yet poorly controlled variable in language-guided embodied AI. Existing benchmarks typically pair each task with a single static instruction…
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
Wanjun Du, Zifeng Yuan, Tingting Chen +3
Existing vision-language models (VLMs) have demonstrated impressive performance in reasoning-based segmentation. However, current benchmarks are primarily constructed from high-qua…
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
Fucai Ke, Zhixi Cai, Boying Li +6
Multi-view visual reasoning is essential for intelligent systems that must understand complex environments from sparse and discrete viewpoints, yet existing research has largely fo…