24 citations · 81 across the 62 of their papers we have counts for
7 papers · 2 filters
AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation
Yangchao Wu, Tian Yu Liu, Hyoungseob Park +3
Unsupervised depth completion and estimation methods are trained by minimizing reconstruction error. Block artifacts from resampling, intensity saturation, and occlusions are among…
Sub-token ViT Embedding via Stochastic Resonance Transformers
Dong Lao, Yangchao Wu, Tian Yu Liu +2
Vision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This repres…
Towards Visual Foundational Models of Physical Scenes
Chethan Parameshwara, Alessandro Achille, Matthew Trager +7
We describe a first step towards learning general-purpose visual representations of physical scenes using only image prediction as a training criterion. To do so, we first define "…
Prompt Algebra for Task Composition
Pramuditha Perera, Matthew Trager, Luca Zancato +2
We investigate whether prompts learned independently for different tasks can be later combined through prompt algebra to obtain a model that supports composition of tasks. We consi…
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
Zhaoyang Zhang, Yantao Shen, Kunyu Shi +7
We present a vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple heterogeneous tasks which may interfere with each other, result…
Train/Test-Time Adaptation with Retrieval
Luca Zancato, Alessandro Achille, Tian Yu Liu +3
We introduce Train/Test-Time Adaptation with Retrieval (), a method to adapt models both at train and test time by means of a retrieval module and a searchable pool of…