Visual Intelligence through Human Interaction
arXiv:2111.06913 · doi:10.1007/978-3-030-82681-9
Abstract
Over the last decade, Computer Vision, the branch of Artificial Intelligence aimed at understanding the visual world, has evolved from simply recognizing objects in images to describing pictures, answering questions about images, aiding robots maneuver around physical spaces and even generating novel visual content. As these tasks and applications have modernized, so too has the reliance on more data, either for model training or for evaluation. In this chapter, we demonstrate that novel interaction strategies can enable new forms of data collection and evaluation for Computer Vision. First, we present a crowdsourcing interface for speeding up paid data collection by an order of magnitude, feeding the data-hungry nature of modern vision models. Second, we explore a method to increase volunteer contributions using automated social interventions. Third, we develop a system to ensure human evaluation of generative vision models are reliable, affordable and grounded in psychophysics theory. We conclude with future opportunities for Human-Computer Interaction to aid Computer Vision.
This is a preprint of the following chapter: Ranjay Krishna, Mitchell Gordon, Li Fei-Fei, Michael Bernstein, Visual Intelligence through Human Interaction, published in Artificial Intelligence for Human Computer Interaction: A Modern Approach, edited by Yang Li and Otmar Hilliges, 2021, Springer reproduced with permission of Springer Nature. arXiv admin note: substantial text overlap with arXiv:1602.04506, arXiv:1904.01121
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- A Neural Representation of Sketch Drawings
- Variational Approaches for Auto-Encoding Generative Adversarial Networks
- Conceptual Metaphors Impact Perceptions of Human-AI Collaboration
- Evorus: A Crowd-powered Conversational Assistant Built to Automate Itself Over Time
- Cross-Modal Hierarchical Modelling for Fine-Grained Sketch Based Image Retrieval
- Iris: A Conversational Agent for Complex Tasks
- CoSE: Compositional Stroke Embeddings
- The DIDI dataset: Digital Ink Diagram data
- Sketchforme: Composing Sketched Scenes from Text Descriptions for Interactive Applications
- DeepWriting: Making Digital Ink Editable via Deep Generative Modeling
- How2Sketch: Generating Easy-To-Follow Tutorials for Sketching 3D Objects