activity
20232026
most citedPOP-3D: Open-Vocabulary 3D Occupancy Prediction from Images

8 citations · 18 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Autoregressive Modeling of Film with Applications in Video Montage

Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar +4

This work introduces FilmGPT, an autoregressive transformer designed to address the challenge of video montage--turning a collection of raw, "unwatchable" footage into coherent cin…

cs.CV2025

AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment

Anna Šárová Mikeštíková, Médéric Fourmy, Martin Cífka +2

Single-view RGB model-based object pose estimation methods achieve strong generalization but are fundamentally limited by depth ambiguity, clutter, and occlusions. Multi-view pose…

cs.CV2025

ResidualViT for Efficient Temporally Dense Video Encoding

Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem +2

Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" r…

cs.CV20256 cited

EditDuet: A Multi-Agent System for Video Non-Linear Editing

Marcelo Sandoval-Castaneda, Bryan Russell, Josef Sivic +2

Automated tools for video editing and assembly have applications ranging from filmmaking and advertisement to content creation for social media. Previous video editing work has mai…

cs.CV2025

Discovering Divergent Representations between Text-to-Image Models

Lisa Dunlap, Joseph E. Gonzalez, Trevor Darrell +3

In this paper, we investigate when and how visual representations learned by two different generative models diverge. Given two text-to-image models, our goal is to discover visual…

cs.CV2025

Improving Personalized Search with Regularized Low-Rank Parameter Updates

Fiona Ryan, Josef Sivic, Fabian Caba Heilbron +3

Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning…