activity
20142024
most citedSupervised Descent Method for Solving Nonlinear Least Squares Problems in Computer Vision

44 citations · 57 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

Streaming Dense Video Captioning

Xingyi Zhou, Anurag Arnab, Shyamal Buch +5

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descr…

cs.CV20232 cited

UnLoc: A Unified Framework for Video Localization Tasks

Shen Yan, Xuehan Xiong, Arsha Nagrani +5

While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos…

cs.CV20233 cited

End-to-End Spatio-Temporal Action Localisation with Video Transformers

Alexey Gritsenko, Xuehan Xiong, Josip Djolonga +5

The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end, purely-transformer ba…

cs.CV20224 cited

Beyond Transfer Learning: Co-finetuning for Action Localisation

Anurag Arnab, Xuehan Xiong, Alexey Gritsenko +6

Transfer learning is the predominant paradigm for training deep networks on small target datasets. Models are typically pretrained on large ``upstream'' datasets for classification…

cs.CV20224 cited

Multiview Transformers for Video Recognition

Shen Yan, Xuehan Xiong, Anurag Arnab +4

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer…

cs.CV201444 cited

Supervised Descent Method for Solving Nonlinear Least Squares Problems in Computer Vision

Xuehan Xiong, Fernando De la Torre

Many computer vision problems (e.g., camera calibration, image alignment, structure from motion) are solved with nonlinear optimization methods. It is generally accepted that secon…