activity
20192025
most citedA3D: Adaptive 3D Networks for Video Action Recognition

10 citations · 36 across the 7 of their papers we have counts for

collaborators

12 papers

cs.CV2025

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

Andong Deng, Taojiannan Yang, Shoubin Yu +5

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and ch…

cs.CV2024

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

Andong Deng, Tongjia Chen, Shoubin Yu +6

In this paper, we introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the inp…

cs.CV20221 cited

Revisiting Training-free NAS Metrics: An Efficient Training-based Method

Taojiannan Yang, Linjie Yang, Xiaojie Jin +1

Recent neural architecture search (NAS) works proposed training-free metrics to rank networks which largely reduced the search cost in NAS. In this paper, we revisit these training…

cs.CV2021

3D Human Pose Estimation with Spatial and Temporal Transformers

Ce Zheng, Sijie Zhu, Matias Mendieta +3

Transformer architectures have become the model of choice in natural language processing and are now being introduced into computer vision tasks such as image classification, objec…

cs.CV202010 cited

A3D: Adaptive 3D Networks for Video Action Recognition

Sijie Zhu, Taojiannan Yang, Matias Mendieta +1

This paper presents A3D, an adaptive 3D network that can infer at a wide range of computational constraints with one-time training. Instead of training multiple models in a grid-se…

cs.CV20208 cited

Cross-directional Feature Fusion Network for Building Damage Assessment from Satellite Imagery

Yu Shen, Sijie Zhu, Taojiannan Yang +1

Fast and effective responses are required when a natural disaster (e.g., earthquake, hurricane, etc.) strikes. Building damage assessment from satellite imagery is critical before…