papers

Publications (30)

cs.CV2021

Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

Reuben Tan, Bryan A. Plummer, Kate Saenko +2

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision…

cs.CV2019

Language Features Matter: Effective Language Representations for Vision-Language Tasks

Andrea Burns, Reuben Tan, Kate Saenko +2

Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language m…

cs.CV2025

Magma: A Foundation Model for Multimodal AI Agents

Jianwei Yang, Reuben Tan, Qianhui Wu +10

We present Magma, a foundation model that serves multimodal AI agentic tasks in both the digital and physical worlds. Magma is a significant extension of vision-language (VL) model…

cs.CV2024

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Mu Cai, Reuben Tan, Jianrui Zhang +12

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benc…

cs.AI2026

Building Agent Harnesses for Scientific Curation from Multimodal Sources

Sheng Zhang, Qin Liu, Renqian Luo +9

The paper introduces Beaver, an agent harness that extracts structured scientific information from papers by integrating text, tables, and figures while preserving provenance, and…

#scientific curation#multimodal information extraction#provenance tracking#agent harness design
cs.CV2025

MindJourney: Test-Time Scaling with World Models for Spatial Reasoning

Yuncong Yang, Jiageng Liu, Zheyuan Zhang +5

Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision-language mode…

cs.CV2022

NewsStories: Illustrating articles with visual summaries

Reuben Tan, Bryan A. Plummer, Kate Saenko +3

Recent self-supervised approaches have used large-scale image-text datasets to learn powerful representations that transfer to many tasks without finetuning. These methods often as…

cs.AI2026

AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback

Andrea Tupini, Lars Liden, Reuben Tan +2

With AsgardBench we aim to evaluate visually grounded, high-level action sequence generation and interactive planning, focusing specifically on plan adaptation during execution bas…

cs.CV2023

Multiscale Video Pretraining for Long-Term Activity Forecasting

Reuben Tan, Matthias De Lange, Michael Iuzzolino +4

Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the v…

cs.CV2020

LoGAN: Latent Graph Co-Attention Network for Weakly-Supervised Video Moment Retrieval

Reuben Tan, Huijuan Xu, Kate Saenko +1

The goal of weakly-supervised video moment retrieval is to localize the video segment most relevant to the given natural language query without access to temporal annotations durin…

cs.CV2026

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Yao Xiao, Reuben Tan, Zhen Zhu +3

ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…

#vision-language models#visual retrieval#sparse token selection#long video processing
cs.RO2026

Spatially Grounded Long-Horizon Task Planning in the Wild

Sehun Jung, HyunJee Song, Dong-Hee Kim +4

Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action pla…

cs.CV2023

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

Reuben Tan, Arijit Ray, Andrea Burns +5

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as tra…

cs.AI2020

Detecting Cross-Modal Inconsistency to Defend Against Neural Fake News

Reuben Tan, Bryan A. Plummer, Kate Saenko

Large-scale dissemination of disinformation online intended to mislead or deceive the general population is a major societal problem. Rapid progression in image, video, and natural…

cs.AI2026

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Reuben Tan, Baolin Peng, Zhengyuan Yang +16

Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-bas…

cs.CL2025

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Qianhui Wu, Kanzhi Cheng, Rui Yang +15

One of the principal challenges in building VLM-powered GUI agents is visual grounding, i.e., localizing the appropriate screen region for action execution based on both the visual…

cs.AI2025

The Illusion of Readiness in Health AI

Yu Gu, Jingjing Fu, Xiaodong Liu +29

Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, espec…

cs.AI2025

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

Timothy Ossowski, Sheng Zhang, Qianchu Liu +5

High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tas…

cs.CV2024

Koala: Key frame-conditioned long video-LLM

Reuben Tan, Ximeng Sun, Ping Hu +5

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Lar…

cs.CV2026

Learning Sparse Visual Representations via Spatial-Semantic Factorization

Theodore Zhengde Zhao, Sid Kiblawi, Jianwei Yang +6

Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens th…

cs.CV2026

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

Dong-Hee Kim, Reuben Tan, Donghyun Kim

Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search ta…

cs.CV2026

InstrAct: Towards Action-Centric Understanding in Instructional Videos

Zhuoyi Yang, Jiapeng Yu, Reuben Tan +2

Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (V…

cs.RO2025

Latent Action Pretraining from Videos

Seonghyeon Ye, Joel Jang, Byeongguk Jeon +13

We introduce Latent Action Pretraining for general Action models (LAPA), an unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot actio…

cs.LG2025

Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs

Aoming Liu, Reuben Tan, Boqing Gong +1

Token reduction accelerates Multimodal Large Language Models (MLLMs) by reducing excessive tokens, but overlooks structural redundancy differences, where critical and redundant mod…

cs.AI2023

Socratis: Are large multimodal models emotionally aware?

Katherine Deng, Arijit Ray, Reuben Tan +3

Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reas…

cs.CV2026

VideoWeave: A Data-Centric Approach for Efficient Video Understanding

Zane Durante, Silky Singh, Arpandeep Khatua +6

Training video-language models is often prohibitively expensive due to the high cost of processing long frame sequences and the limited availability of annotated long videos. We pr…

cs.CV2023

EgoAdapt: A multi-stream evaluation study of adaptation to real-world egocentric user video

Matthias De Lange, Hamid Eghbalzadeh, Reuben Tan +3

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this…

cs.CV2019

Learning Similarity Conditions Without Explicit Supervision

Reuben Tan, Mariya I. Vasileva, Kate Saenko +1

Many real-world tasks require models to compare images along multiple similarity conditions (e.g. similarity in color, category or shape). Existing methods often reason about these…

cs.CV2025

SITE: towards Spatial Intelligence Thorough Evaluation

Wenqi Wang, Reuben Tan, Pengyue Zhu +6

Spatial intelligence (SI) represents a cognitive ability encompassing the visualization, manipulation, and reasoning about spatial relationships, underpinning disciplines from neur…

cs.CV2025

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

Arijit Ray, Jiafei Duan, Ellis Brown +9

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal lang…