Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Long Cheng, Jiafei Duan, Yi Ru Wang +12
Pointing serves as a fundamental and intuitive mechanism for grounding language within visual contexts, with applications spanning robotics, assistive technologies, and interactive…
cs.CV2021
Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans
Ainaz Eftekhar, Alexander Sax, Roman Bachmann +2
This paper introduces a pipeline to parametrically sample and render multi-task vision datasets from comprehensive 3D scans from the real world. Changing the sampling parameters al…