Publications (22)
Non-Markov Policies to Reduce Sequential Failures in Robot Bin Picking
Kate Sanders, Michael Danielczuk, Jeffrey Mahler +2
A new generation of automated bin picking systems using deep learning is evolving to support increasing demand for e-commerce. To accommodate a wide variety of products, many autom…
A Survey of Video Datasets for Grounded Event Understanding
Kate Sanders, Benjamin Van Durme
While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-…
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
Kate Sanders, Benjamin Van Durme
To develop general-purpose collaborative agents, humans need reliable AI systems that can (1) adapt to new domains and (2) transparently reason with uncertainty to allow for verifi…
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
Kate Sanders, Nathaniel Weir, Sapana Chaudhary +2
An impediment to using Large Language Models (LLMs) for reasoning output verification is that LLMs struggle to reliably identify errors in thinking traces, particularly in long out…
Ambiguous Images With Human Judgments for Robust Visual Event Classification
Kate Sanders, Reno Kriz, Anqi Liu +1
Contemporary vision benchmarks predominantly consider tasks on which humans can achieve near-perfect performance. However, humans are frequently presented with visual data that the…
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
Jiefu Ou, William Gantt Walden, Kate Sanders +13
A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically genera…
Grounding Partially-Defined Events in Multimodal Data
Kate Sanders, Reno Kriz, David Etter +5
How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially…
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
Alexander Martin, William Walden, Reno Kriz +5
We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a prevalent source of information online…
WikiVideo: Article Generation from Multiple Videos
Alexander Martin, Reno Kriz, William Gantt Walden +5
We introduce the task of grounded article generation with the goal of creating a Wikipedia-style article from multiple diverse videos about real-world events -- from natural disast…
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
Nathaniel Weir, Kate Sanders, Orion Weller +8
Recent language models enable new opportunities for structured reasoning with text, such as the construction of intuitive, proof-like textual entailment trees without relying on br…
A Multi-Chamber Smart Suction Cup for Adaptive Gripping and Haptic Exploration
Tae Myung Huh, Kate Sanders, Michael Danielczuk +4
We present a novel robot end-effector for gripping and haptic exploration. Tactile sensing through suction flow monitoring is applied to a new suction cup design that contains mult…
Mechanical Search on Shelves using Lateral Access X-RAY
Huang Huang, Marcus Dominguez-Kuhne, Jeffrey Ichnowski +7
Efficiently finding an occluded object with lateral access arises in many contexts such as warehouses, retail, healthcare, shipping, and homes. We introduce LAX-RAY (Lateral Access…
Core: Robust Factual Precision with Informative Sub-Claim Identification
Zhengping Jiang, Jingyu Zhang, Nathaniel Weir +6
Hallucinations pose a challenge to the application of large language models (LLMs) thereby motivating the development of metrics to evaluate factual precision. We observe that popu…
MultiVENT: Multilingual Videos of Events with Aligned Natural Text
Kate Sanders, David Etter, Reno Kriz +1
Everyday news coverage has shifted from traditional broadcasts towards a wide range of presentation formats such as first-hand, unedited video footage. Datasets that reflect the di…
MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion
Saron Samuel, Dan DeGenaro, Jimena Guallar-Blasco +13
Videos inherently contain multiple modalities, including visual events, text overlays, sounds, and speech, all of which are important for retrieval. However, state-of-the-art multi…
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
Reno Kriz, Kate Sanders, David Etter +10
Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from…
On the Evaluation of Machine-Generated Reports
James Mayfield, Eugene Yang, Dawn Lawrie +10
Large Language Models (LLMs) have enabled new ways to satisfy information needs. Although great strides have been made in applying them to settings like document ranking and short-…
Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning
Kavi Gupta, Kate Sanders, Armando Solar-Lezama
While LLMs have revolutionized the field of machine learning due to their high performance on a strikingly wide range of problems, they are also known to hallucinate false answers…
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
Arun Reddy, Alexander Martin, Eugene Yang +7
In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval…
Tur[k]ingBench: A Challenge Benchmark for Web Agents
Kevin Xu, Yeganeh Kordi, Tanay Nayak +7
Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-ta…
SocialNLI: A Dialogue-Centric Social Inference Dataset
Akhil Deo, Kate Sanders, Benjamin Van Durme
Making theory-of-mind inferences from human dialogue is a strong indicator of a model's underlying social abilities, which are fundamental for adept AI assistants. However, large l…
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
Kate Sanders, Nathaniel Weir, Benjamin Van Durme
It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality rea…