Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
Rajeev Yasarla, Manish Kumar Singh, Hong Cai +6
In this paper, we propose a novel video depth estimation approach, FutureDepth, which enables the model to implicitly leverage multi-frame and motion cues to improve depth estimati…
cs.CV2024
PADRe: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision Transformer
Pierre-David Letourneau, Manish Kumar Singh, Hsin-Pai Cheng +6
We present Polynomial Attention Drop-in Replacement (PADRe), a novel and unifying framework designed to replace the conventional self-attention mechanism in transformer models. Not…
cs.CV2024
ToSA: Token Selective Attention for Efficient Vision Transformers
Manish Kumar Singh, Rajeev Yasarla, Hong Cai +2
In this paper, we propose a novel token selective attention approach, ToSA, which can identify tokens that need to be attended as well as those that can skip a transformer layer. M…