118 citations · 118 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 118 cited
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
cs.CV2023
Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features
Vivek Rathod, Bryan Seybold, Sudheendra Vijayanarasimhan +4
Detecting actions in untrimmed videos should not be limited to a small, closed set of classes. We present a simple, yet effective strategy for open-vocabulary temporal action detec…