Publications (33)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Christopher Clark, Yue Yang, Jae Sung Park +8
Grounding has become a fundamental capability of vision-language models (VLMs). Most existing VLMs point by generating coordinates as part of their text output, which requires lear…
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
Jianrui Zhang, Yue Yang, Rohun Tripathi +5
Token pruning is essential for enhancing the computational efficiency of vision-language models (VLMs), particularly for video-based tasks where temporal redundancy is prevalent. P…
Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Jiasen Lu, Christopher Clark, Rowan Zellers +2
We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and…
Holodeck: Language Guided Generation of 3D Embodied AI Environments
Yue Yang, Fan-Yun Sun, Luca Weihs +11
3D simulated environments play a critical role in Embodied AI, but their creation requires expertise and extensive manual effort, restricting their diversity and scope. To mitigate…
The Causes of the Red Sequence, the Blue Cloud, the Green Valley and the Green Mountain
Stephen Eales, Maarten Baes, Nathan Bourne +23
The galaxies found in optical surveys fall in two distinct regions of a diagram of optical colour versus absolute magnitude: the red sequence and the blue cloud with the green vall…
Astro2020: Unleashing the Potential of Dust Emission as a Window onto Galaxy Evolution
Christopher Clark, Julua Roman-Duval, Sarah Sadavoy +15
We present the severe, systematic uncertainties currently facing our understanding of dust emission, which stymie our ability to truly exploit dust as a tool for studying galaxy ev…
A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark +2
The Visual Question Answering (VQA) task aspires to provide a meaningful testbed for the development of AI models that can jointly reason over visual and natural language inputs. D…
Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text
Christopher Clark, Jordi Salvador, Dustin Schwenk +13
Communicating with humans is challenging for AIs because it requires a shared understanding of the world, complex semantics (e.g., metaphors or analogies), and at times multi-modal…
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
Chris Dongjoo Kim, Jihwan Moon, Sangwoo Moon +7
The rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising…
The Astropy Problem
Demitri Muna, Michael Alexander, Alice Allen +151
The Astropy Project (http://astropy.org) is, in its own words, "a community effort to develop a single core package for Astronomy in Python and foster interoperability between Pyth…
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Matt Deitke, Christopher Clark, Sangho Lee +47
Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good perfor…
Webly Supervised Concept Expansion for General Purpose Vision Models
Amita Kamath, Christopher Clark, Tanmay Gupta +3
General Purpose Vision (GPV) systems are models that are designed to solve a wide array of visual tasks without requiring architectural changes. Today, GPVs primarily learn both sk…
Learning to Model and Ignore Dataset Bias with Mixed Capacity Ensembles
Christopher Clark, Mark Yatskar, Luke Zettlemoyer
Many datasets have been shown to contain incidental correlations created by idiosyncrasies in the data collection process. For example, sentence entailment datasets can have spurio…
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Jiasen Lu, Christopher Clark, Sangho Lee +5
We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we…
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Christopher Clark, Jieyu Zhang, Zixian Ma +18
Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from…
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
Adyasha Maharana, Amita Kamath, Christopher Clark +2
As general purpose vision models get increasingly effective at a wide set of tasks, it is imperative that they be consistent across the tasks they support. Inconsistent AI models a…
The PRIMA promise of deciphering interstellar dust evolution with observations of the nearby Universe
Frédéric Galliano, Maarten Baes, Léo Belloir +25
This paper develops a few science cases, using the PRIMA far-IR probe, aimed at achieving several breakthroughs in our understanding of the dust properties and their evolution. We…
Interstellar Dust Grains: Ultraviolet and Mid-IR Extinction Curves
Karl D. Gordon, Karl Misselt, Yvonne Pendleton +7
Interstellar dust plays a central role in shaping the detailed structure of the interstellar medium, thus strongly influencing star formation and galaxy evolution. Dust extinction…
Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases
Christopher Clark, Mark Yatskar, Luke Zettlemoyer
State-of-the-art models often make use of superficial patterns in the data that do not generalize well to out-of-domain or adversarial settings. For example, textual entailment mod…
I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision
Sophia Gu, Christopher Clark, Aniruddha Kembhavi
Many high-level skills that are required for computer vision tasks, such as parsing questions, comparing and contrasting semantics, and writing descriptions, are also required in o…
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer +4
We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.g., syntax and semantics), and (2) how these uses var…
Teaching Deep Convolutional Neural Networks to Play Go
Christopher Clark, Amos Storkey
Mastering the game of Go has remained a long standing challenge to the field of AI. Modern computer Go systems rely on processing millions of possible future positions to play well…
The Life Cycle of Dust
Sarah Sadavoy, Mikako Matsuura, Lee Armus +24
Dust offers a unique probe of the interstellar medium (ISM) across multiple size, density, and temperature scales. Dust is detected in outflows of evolved stars, star-forming molec…
Multi-AUV Marine Life Tracking with Single Hydrophone Payloads via a Hidden Markov Model Equipped Particle Filter
Christopher Herrera, Kehlani Fay, Christopher Clark +3
Researchers tag and track marine animals to study migration patterns, human impacts on behavior, and behavioral shifts due to climate change. Accurate data collection often require…
Classification for Big Dataset of Bioacoustic Signals Based on Human Scoring System and Artificial Neural Network
Mohammad Pourhomayoun, Peter Dugan, Marian Popescu +3
In this paper, we propose a method to improve sound classification performance by combining signal features, derived from the time-frequency spectrogram, with human perception. The…
Long-distance Detection of Bioacoustic Events with Per-channel Energy Normalization
Vincent Lostanlen, Kaitlin Palmer, Elly Knight +6
This paper proposes to perform unsupervised detection of bioacoustic events by pooling the magnitudes of spectrogram frames after per-channel energy normalization (PCEN). Although…
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
Jitesh Jain, Jialuo Li, Zixian Ma +7
As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this…
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Christopher Clark, Kenton Lee, Ming-Wei Chang +3
In this paper we study yes/no questions that are naturally occurring --- meaning that they are generated in unprompted and unconstrained settings. We build a reading comprehension…
One Diffusion to Generate Them All
Duong H. Le, Tuan Pham, Sangho Lee +5
We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables condit…
2 OLMo 2 Furious
Team OLMo, Pete Walsh, Luca Soldaini +40
We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully rele…
Simple and Effective Multi-Paragraph Reading Comprehension
Christopher Clark, Matt Gardner
We consider the problem of adapting neural paragraph-level question answering models to the case where entire documents are given as input. Our proposed solution trains models to p…
Bioacoustic Signal Classification Based on Continuous Region Processing, Grid Masking and Artificial Neural Network
Mohammad Pourhomayoun, Peter Dugan, Marian Popescu +1
In this paper, we develop a novel method based on machine-learning and image processing to identify North Atlantic right whale (NARW) up-calls in the presence of high levels of amb…
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
Yue Yang, Ajay Patel, Matt Deitke +8
Reasoning about images with rich text, such as charts and documents, is a critical application of vision-language models (VLMs). However, VLMs often struggle in these domains due t…