most citedMMViT: Multiscale Multiview Vision Transformers

8 citations · 10 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CL20236 cited

PsyCoT: Psychological Questionnaire as Powerful Chain-of-Thought for Personality Detection

Tao Yang, Tianyuan Shi, Fanqi Wan +4

Recent advances in large language models (LLMs), such as ChatGPT, have showcased remarkable zero-shot performance across various NLP tasks. However, the potential of LLMs in person…

cs.CV20237 cited

ClusterFormer: Clustering As A Universal Visual Learner

James C. Liang, Yiming Cui, Qifan Wang +3

This paper presents CLUSTERFORMER, a universal vision model that is based on the CLUSTERing paradigm with TransFORMER. It comprises two novel designs: 1. recurrent cross-attention…

cs.CV20238 cited

MMViT: Multiscale Multiview Vision Transformers

Yuchen Liu, Natasha Ong, Kaiyan Peng +8

We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…

cs.CV20231 cited

TransFlow: Transformer as Flow Learner

Yawen Lu, Qifan Wang, Siqi Ma +4

Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work,…

cs.CV2023

SVT: Supertoken Video Transformer for Efficient Video Understanding

Chenbin Pan, Rui Hou, Hanchao Yu +3

Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…

cs.SD20221 cited

Fall Detection from Audios with Audio Transformers

Prabhjot Kaur, Qifan Wang, Weisong Shi

Fall detection for the elderly is a well-researched problem with several proposed solutions, including wearable and non-wearable techniques. While the existing techniques have exce…