4 papers
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation
Yifan Yin, Zhengtao Han, Shivam Aarya +6
Fine-grained robot manipulation, such as lifting and rotating a bottle to display the label on the cap, requires robust reasoning about object parts and their relationships with in…
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
Kui Wu, Shuhang Xu, Hao Chen +4
We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tr…
CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language Games
Shuhang Xu, Fangwei Zhong
Metaphors are a crucial way for humans to express complex or subtle ideas by comparing one concept to another, often from a different domain. However, many large language models (L…
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
Shuhang Xu, Weijian Deng, Yixuan Zhou +1
Concepts serve as fundamental abstractions that support human reasoning and categorization. However, it remains unclear whether large language models truly capture such conceptual…