activity
20202026
most citedRevisiting 3D ResNets for Video Recognition

15 citations · 36 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models

Yunkai Zhang, Linda Li, Yingxin Cui +5

Vision-Language Models (VLMs) excel on many multimodal reasoning benchmarks, but these evaluations often do not require an exhaustive readout of the image and can therefore obscure…

cs.CV2023

VideoGLUE: Video General Understanding Evaluation of Foundation Models

Liangzhe Yuan, Nitesh Bharadwaj Gundavarapu, Long Zhao +14

We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recog…

cs.CV2023

Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception

Hassan Akbari, Dan Kondratyuk, Yin Cui +3

We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, vid…

cs.CV2023

Unified Visual Relationship Detection with Vision and Language Models

Long Zhao, Liangzhe Yuan, Boqing Gong +5

This work focuses on training a single visual relationship detector predicting over the union of label spaces from multiple datasets. Merging labels spanning different datasets cou…

cs.CV202115 cited

Revisiting 3D ResNets for Video Recognition

Xianzhi Du, Yeqing Li, Yin Cui +3

A recent work from Bello shows that training and scaling strategies may be more significant than model architectures for visual recognition. This short note studies effective train…

cs.CV20214 cited

Federated Multi-Target Domain Adaptation

Chun-Han Yao, Boqing Gong, Yin Cui +3

Federated learning methods enable us to train machine learning models on distributed user data while preserving its privacy. However, it is not always feasible to obtain high-quali…