64 citations · 84 across the 16 of their papers we have counts for
29 papers
Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations
Denis M. Akola, David F. Fouhey
3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene repre…
Visual General Intelligence: A White Paper
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…
The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning
V. Samuel Pérez-Díaz, Vinay L. Kashyap, Joshua D. Ingram +6
We present a framework to cross-match sources from the Chandra Source Catalog (CSC v2.1) with optical sources from Gaia Data Release 3. Unlike purely spatial approaches, we use sou…
Improving and Evaluating Hand-Object Interaction Detection
Ahmad Darkhalil, Dima Damen, David Fouhey
Understanding hands and the objects they interact with, both directly and through tools, is a key step for tasks ranging from action perception to 3D reconstruction and robotics. O…
Human Universal Grasping
Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu +5
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from hum…
Dynamic Camera Poses and Where to Find Them
Chris Rockwell, Joseph Tung, Tsung-Yi Lin +3
Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is d…