From the 1 of 7 linked papers with an AI index.
7 papers
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
The paper investigates whether large language models (LLMs) have distinct planning abilities by applying multidimensional item response theory to benchmark data, uncovering two sep…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fucai Ke, Joy Hsu, Zhixi Cai +10
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…
Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents
Sukai Huang, Chenyuan Zhang, Fucai Ke +4
Instruction granularity is an important yet poorly controlled variable in language-guided embodied AI. Existing benchmarks typically pair each task with a single static instruction…
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
Zhixi Cai, Fucai Ke, Kevin Leo +4
Recent vision-language models have strong perceptual ability but their implicit reasoning is hard to explain and easily generates hallucinations on complex queries. Compositional m…
ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
Chunhua Liu, Kabir Manandhar Shrestha, Sukai Huang
Large language models (LLMs) exhibit cultural bias from overrepresented viewpoints in training data, yet cultural alignment remains a challenge due to limited cultural knowledge an…
The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
Sukai Huang, Shu-Wei Liu, Nir Lipovetzky +1
While Vision-Language Models (VLMs) are increasingly used to generate reward signals for training embodied agents to follow instructions, our research reveals that agents guided by…