papers

Publications (33)

cs.CV2026

MolmoPoint: Better Pointing for VLMs with Grounding Tokens

Christopher Clark, Yue Yang, Jae Sung Park +8

Grounding has become a fundamental capability of vision-language models (VLMs). Most existing VLMs point by generating coordinates as part of their text output, which requires lear…

cs.CV2026

Unified Spatio-Temporal Token Scoring for Efficient Video VLMs

Jianrui Zhang, Yue Yang, Rohun Tripathi +5

Token pruning is essential for enhancing the computational efficiency of vision-language models (VLMs), particularly for video-based tasks where temporal redundancy is prevalent. P…

cs.CV2022

Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Jiasen Lu, Christopher Clark, Rowan Zellers +2

We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and…

cs.CV2024

Holodeck: Language Guided Generation of 3D Embodied AI Environments

Yue Yang, Fan-Yun Sun, Luca Weihs +11

3D simulated environments play a critical role in Embodied AI, but their creation requires expertise and extensive manual effort, restricting their diversity and scope. To mitigate…

astro-ph.GA2018

The Causes of the Red Sequence, the Blue Cloud, the Green Valley and the Green Mountain

Stephen Eales, Maarten Baes, Nathan Bourne +23

The galaxies found in optical surveys fall in two distinct regions of a diagram of optical colour versus absolute magnitude: the red sequence and the blue cloud with the green vall…

astro-ph.GA2019

Astro2020: Unleashing the Potential of Dust Emission as a Window onto Galaxy Evolution

Christopher Clark, Julua Roman-Duval, Sarah Sadavoy +15

We present the severe, systematic uncertainties currently facing our understanding of dust emission, which stymie our ability to truly exploit dust as a tool for studying galaxy ev…

cs.CV2022

A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Dustin Schwenk, Apoorv Khandelwal, Christopher Clark +2

The Visual Question Answering (VQA) task aspires to provide a meaningful testbed for the development of AI models that can jointly reason over visual and natural language inputs. D…

cs.CL2021

Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text

Christopher Clark, Jordi Salvador, Dustin Schwenk +13

Communicating with humans is challenging for AIs because it requires a shared understanding of the world, complex semantics (e.g., metaphors or analogies), and at times multi-modal…

cs.CV2025

ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams

Chris Dongjoo Kim, Jihwan Moon, Sangwoo Moon +7

The rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising…

astro-ph.IM2016

The Astropy Problem

Demitri Muna, Michael Alexander, Alice Allen +151

The Astropy Project (http://astropy.org) is, in its own words, "a community effort to develop a single core package for Astronomy in Python and foster interoperability between Pyth…

cs.CV2024

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Matt Deitke, Christopher Clark, Sangho Lee +47

Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good perfor…

cs.CV2022

Webly Supervised Concept Expansion for General Purpose Vision Models

Amita Kamath, Christopher Clark, Tanmay Gupta +3

General Purpose Vision (GPV) systems are models that are designed to solve a wide array of visual tasks without requiring architectural changes. Today, GPVs primarily learn both sk…

cs.LG2020

Learning to Model and Ignore Dataset Bias with Mixed Capacity Ensembles

Christopher Clark, Mark Yatskar, Luke Zettlemoyer

Many datasets have been shown to contain incidental correlations created by idiosyncrasies in the data collection process. For example, sentence entailment datasets can have spurio…

cs.CV2023

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Jiasen Lu, Christopher Clark, Sangho Lee +5

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we…

cs.CV2026

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Christopher Clark, Jieyu Zhang, Zixian Ma +18

Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from…

cs.CV2024

Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models

Adyasha Maharana, Amita Kamath, Christopher Clark +2

As general purpose vision models get increasingly effective at a wide set of tasks, it is imperative that they be consistent across the tasks they support. Inconsistent AI models a…

astro-ph.IM2025

The PRIMA promise of deciphering interstellar dust evolution with observations of the nearby Universe

Frédéric Galliano, Maarten Baes, Léo Belloir +25

This paper develops a few science cases, using the PRIMA far-IR probe, aimed at achieving several breakthroughs in our understanding of the dust properties and their evolution. We…

astro-ph.GA2019

Interstellar Dust Grains: Ultraviolet and Mid-IR Extinction Curves

Karl D. Gordon, Karl Misselt, Yvonne Pendleton +7

Interstellar dust plays a central role in shaping the detailed structure of the interstellar medium, thus strongly influencing star formation and galaxy evolution. Dust extinction…

cs.CL2019

Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases

Christopher Clark, Mark Yatskar, Luke Zettlemoyer

State-of-the-art models often make use of superficial patterns in the data that do not generalize well to out-of-domain or adversarial settings. For example, textual entailment mod…

cs.CV2023

I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision

Sophia Gu, Christopher Clark, Aniruddha Kembhavi

Many high-level skills that are required for computer vision tasks, such as parsing questions, comparing and contrasting semantics, and writing descriptions, are also required in o…

cs.CL2018

Deep contextualized word representations

Matthew E. Peters, Mark Neumann, Mohit Iyyer +4

We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.g., syntax and semantics), and (2) how these uses var…

cs.AI2015

Teaching Deep Convolutional Neural Networks to Play Go

Christopher Clark, Amos Storkey

Mastering the game of Go has remained a long standing challenge to the field of AI. Modern computer Go systems rely on processing millions of possible future positions to play well…

astro-ph.GA2019

The Life Cycle of Dust

Sarah Sadavoy, Mikako Matsuura, Lee Armus +24

Dust offers a unique probe of the interstellar medium (ISM) across multiple size, density, and temperature scales. Dust is detected in outflows of evolved stars, star-forming molec…

cs.RO2026

Multi-AUV Marine Life Tracking with Single Hydrophone Payloads via a Hidden Markov Model Equipped Particle Filter

Christopher Herrera, Kehlani Fay, Christopher Clark +3

Researchers tag and track marine animals to study migration patterns, human impacts on behavior, and behavioral shifts due to climate change. Accurate data collection often require…

cs.CV2013

Classification for Big Dataset of Bioacoustic Signals Based on Human Scoring System and Artificial Neural Network

Mohammad Pourhomayoun, Peter Dugan, Marian Popescu +3

In this paper, we propose a method to improve sound classification performance by combining signal features, derived from the time-frequency spectrogram, with human perception. The…

cs.SD2019

Long-distance Detection of Bioacoustic Events with Per-channel Energy Normalization

Vincent Lostanlen, Kaitlin Palmer, Elly Knight +6

This paper proposes to perform unsupervised detection of bioacoustic events by pooling the magnitudes of spectrogram frames after per-channel energy normalization (PCEN). Although…

cs.CV2026

SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning

Jitesh Jain, Jialuo Li, Zixian Ma +7

As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this…

cs.CL2019

BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Christopher Clark, Kenton Lee, Ming-Wei Chang +3

In this paper we study yes/no questions that are naturally occurring --- meaning that they are generated in unprompted and unconstrained settings. We build a reading comprehension…

cs.CV2025

One Diffusion to Generate Them All

Duong H. Le, Tuan Pham, Sangho Lee +5

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables condit…

cs.CL2025

2 OLMo 2 Furious

Team OLMo, Pete Walsh, Luca Soldaini +40

We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully rele…

cs.CL2017

Simple and Effective Multi-Paragraph Reading Comprehension

Christopher Clark, Matt Gardner

We consider the problem of adapting neural paragraph-level question answering models to the case where entire documents are given as input. Our proposed solution trains models to p…

cs.CV2013

Bioacoustic Signal Classification Based on Continuous Region Processing, Grid Masking and Artificial Neural Network

Mohammad Pourhomayoun, Peter Dugan, Marian Popescu +1

In this paper, we develop a novel method based on machine-learning and image processing to identify North Atlantic right whale (NARW) up-calls in the presence of high levels of amb…

cs.CV2025

Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

Yue Yang, Ajay Patel, Matt Deitke +8

Reasoning about images with rich text, such as charts and documents, is a critical application of vision-language models (VLMs). However, VLMs often struggle in these domains due t…