activity
20212026
most citedMasked Vision and Language Modeling for Multi-modal Representation Learning

24 citations · 81 across the 62 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

7 papers · 2 filters

cs.CV2023★ 3 cited

AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation

Yangchao Wu, Tian Yu Liu, Hyoungseob Park +3

Unsupervised depth completion and estimation methods are trained by minimizing reconstruction error. Block artifacts from resampling, intensity saturation, and occlusions are among…

cs.CV2023★ 1 cited

Sub-token ViT Embedding via Stochastic Resonance Transformers

Dong Lao, Yangchao Wu, Tian Yu Liu +2

Vision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This repres…

cs.CV2023★ 1 cited

Towards Visual Foundational Models of Physical Scenes

Chethan Parameshwara, Alessandro Achille, Matthew Trager +7

We describe a first step towards learning general-purpose visual representations of physical scenes using only image prediction as a training criterion. To do so, we first define "…

cs.CV2023

Prompt Algebra for Task Composition

Pramuditha Perera, Matthew Trager, Luca Zancato +2

We investigate whether prompts learned independently for different tasks can be later combined through prompt algebra to obtain a model that supports composition of tasks. We consi…

cs.CV2023★ 1 cited

Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts

Zhaoyang Zhang, Yantao Shen, Kunyu Shi +7

We present a vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple heterogeneous tasks which may interfere with each other, result…

cs.CV2023

Train/Test-Time Adaptation with Retrieval

Luca Zancato, Alessandro Achille, Tian Yu Liu +3

We introduce Train/Test-Time Adaptation with Retrieval (), a method to adapt models both at train and test time by means of a retrieval module and a searchable pool of…