collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Visual General Intelligence: A White Paper

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…

cs.CV2026

Beyond Single Object: Learning 3D Relations with Large Language Models

Kohsuke Ide, Ryousuke Yamada, Yue Qiu +4

We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for det…

cs.CV2026

Seeing Red, Thinking Bad: Color Bias in Vision Language Models

Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara +2

Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VL…

cs.CV2025

3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds

Ryousuke Yamada, Kohsuke Ide, Yoshihiro Fukuhara +4

Despite recent progress in 3D self-supervised learning, collecting large-scale 3D scene scans remains expensive and labor-intensive. In this work, we investigate whether 3D represe…

cs.CV2025

Industrial Synthetic Segment Pre-training

Shinichi Mae, Hirokatsu Kataoka, Ryousuke Yamada +3

Vision Foundation Models (VFMs) have made remarkable progress and are increasingly being applied to segmentation tasks in real-world industrial settings. However, VFMs pre-trained…