activity
20242026
collaborators

7 papers

cs.CV2026

Visual General Intelligence: A White Paper

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…

cs.CV2026

Beyond Single Object: Learning 3D Relations with Large Language Models

Kohsuke Ide, Ryousuke Yamada, Yue Qiu +4

We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for det…

cs.CV2026

See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs

Yongchang Zhang, Oliver Ma, Tianyi Liu +2

Recent large vision-language models (LVLMs) have demonstrated impressive reasoning ability by generating long chain-of-thought (CoT) responses. However, CoT reasoning in multimodal…

cs.CL2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

Xianzheng Ma, Tao Sun, Shuai Chen +7

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on te…

cs.CV2025

Inferring Dynamic Physical Properties from Video Foundation Models

Guanqi Zhan, Xianzheng Ma, Weidi Xie +1

We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal information to be inferred: elasti…

cs.RO2025

Robotic Visual Instruction

Yanbang Li, Ziyang Gong, Haoyang Li +4

Recently, natural language has been the primary medium for human-robot interaction. However, its inherent lack of spatial precision introduces challenges for robotic task definitio…