papers

Publications (25)

cs.CV2024

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning

Yicheng Wang, Zhikang Zhang, Jue Wang +6

In various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from tw…

cs.RO2026

Distilling Global Traversability Priors for Image-based Affordance Prediction in Off-road Environments

Matthew Sivaprakasam, Samuel Triest, Micah Nye +5

Standard methods for autonomous navigation in unstructured terrain are prone to myopic behaviors in long-horizon scenarios. The use of metric maps built from LiDAR or cameras provi…

cs.AI2026

Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference

Nenad Banfic, David Fan, Kunal Vaishnavi +5

Deploying high-quality automatic speech recognition (ASR) on edge devices requires models that jointly optimize accuracy, latency, and memory footprint while operating entirely on…

cs.CV2021

Shot Contrastive Self-Supervised Learning for Scene Boundary Detection

Shixing Chen, Xiaohan Nie, David Fan +3

Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boun…

cs.CV2020

OASIS: A Large-Scale Dataset for Single Image 3D in the Wild

Weifeng Chen, Shengyi Qian, David Fan +3

Single-view 3D is the task of recovering 3D properties such as depth and surface normals from a single image. We hypothesize that a major obstacle to single-image 3D is data. We ad…

cs.CV2026

A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures

Basile Terver, Randall Balestriero, Megi Dervishi +8

We present EB-JEPA, an open-source library for learning representations and world models using Joint-Embedding Predictive Architectures (JEPAs). JEPAs learn to predict in represent…

cs.CV2023

Nearest-Neighbor Inter-Intra Contrastive Learning from Unlabeled Videos

David Fan, Deyu Yang, Xinyu Li +2

Contrastive learning has recently narrowed the gap between self-supervised and supervised methods in image and video domain. State-of-the-art video contrastive learning methods suc…

cs.RO2021

Hybrid Imitative Planning with Geometric and Predictive Costs in Off-road Environments

Nitish Dashora, Daniel Shin, Dhruv Shah +5

Geometric methods for solving open-world off-road navigation tasks, by learning occupancy and metric maps, provide good generalization but can be brittle in outdoor environments th…

cs.RO2026

World Models for Learning Dexterous Hand-Object Interactions from Human Videos

Raktim Gautam Goswami, Amir Bar, David Fan +6

Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While rece…

cs.CV2025

Scaling Language-Free Visual Representation Learning

David Fan, Shengbang Tong, Jiachen Zhu +8

Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This mul…

cs.CV2024

Video Token Merging for Long-form Video Understanding

Seon-Ho Lee, Jue Wang, Zhikang Zhang +2

As the scale of data and models for video understanding rapidly expand, handling long-form video input in transformer-based models presents a practical challenge. Rather than resor…

cs.CV2026

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Junlin Han, Shengbang Tong, David Fan +4

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fund…

cs.IR2026

An LLM-powered Agentic Recommendation System for Connected TV Content Discovery

Lei Shi, Di Wang, Harry Tran +22

Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse contextual signals, such as trending topic…

cs.CV2024

NowYouSee Me: Context-Aware Automatic Audio Description

Seon-Ho Lee, Jue Wang, David Fan +5

Audio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable int…

cs.RO2023

A Multi-step Dynamics Modeling Framework For Autonomous Driving In Multiple Environments

Jason Gibson, Bogdan Vlahov, David Fan +4

Modeling dynamics is often the first step to making a vehicle autonomous. While on-road autonomous vehicles have been extensively studied, off-road vehicles pose many challenging m…

cs.CV2024

Text-Guided Video Masked Autoencoder

David Fan, Jue Wang, Shuai Liao +3

Recent video masked autoencoder (MAE) works have designed improved masking algorithms focused on saliency. These works leverage visual cues such as motion to mask the most salient…

cs.CV2023

Motion-Guided Masking for Spatiotemporal Representation Learning

David Fan, Jue Wang, Shuai Liao +5

Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spat…

cs.AI2025

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Mido Assran, Adrien Bardes, David Fan +27

A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-s…

cs.RO2024

RoadRunner -- Learning Traversability Estimation for Autonomous Off-road Driving

Jonas Frey, Manthan Patel, Deegan Atha +7

Autonomous navigation at high speeds in off-road environments necessitates robots to comprehensively understand their surroundings using onboard sensing only. The extreme condition…

cs.RO2022

PrePARE: Predictive Proprioception for Agile Failure Event Detection in Robotic Exploration of Extreme Terrains

Sharmita Dey, David Fan, Robin Schmid +5

Legged robots can traverse a wide variety of terrains, some of which may be challenging for wheeled robots, such as stairs or highly uneven surfaces. However, quadruped robots face…

cs.CV2023

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee +5

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal a…

cond-mat.mtrl-sci2015

Topological defects as relics of emergent continuous symmetry and Higgs condensation of disorder in ferroelectrics

Shi-Zeng Lin, Xueyun Wang, Yoshitomo Kamiya +9

Lars Onsager and Richard Feynman envisioned that the three-dimensional (3D) superfluid-to-normal transition in He occurs through the proliferation of vortices. This proc…

cs.LG2025

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan +4

Large Language Models (LLMs), despite being trained on text alone, surprisingly develop rich visual priors. These priors allow latent visual capabilities to be unlocked for vision…

cs.CV2026

Beyond Language Modeling: An Exploration of Multimodal Pretraining

Shengbang Tong, David Fan, John Nguyen +18

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models r…

cs.CV2024

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Shengbang Tong, David Fan, Jiachen Zhu +7

In this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a pretrained LLM to quickly morph…