7 papers
Visual General Intelligence: A White Paper
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…
Beyond Single Object: Learning 3D Relations with Large Language Models
Kohsuke Ide, Ryousuke Yamada, Yue Qiu +4
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for det…
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
Yongchang Zhang, Oliver Ma, Tianyi Liu +2
Recent large vision-language models (LVLMs) have demonstrated impressive reasoning ability by generating long chain-of-thought (CoT) responses. However, CoT reasoning in multimodal…
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
Xianzheng Ma, Tao Sun, Shuai Chen +7
Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on te…
Inferring Dynamic Physical Properties from Video Foundation Models
Guanqi Zhan, Xianzheng Ma, Weidi Xie +1
We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal information to be inferred: elasti…
Robotic Visual Instruction
Yanbang Li, Ziyang Gong, Haoyang Li +4
Recently, natural language has been the primary medium for human-robot interaction. However, its inherent lack of spatial precision introduces challenges for robotic task definitio…