activity
20212024
most citedVideoLLM: Modeling Video Sequence with Large Language Models

15 citations · 22 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

ActionVOS: Actions as Prompts for Video Object Segmentation

Liangyang Ouyang, Ruicong Liu, Yifei Huang +2

Delving into the realm of egocentric vision, the advancement of referring video object segmentation (RVOS) stands as pivotal in understanding human activities. However, existing RV…

cs.CL2023

Pretraining Language Models with Text-Attributed Heterogeneous Graphs

Tao Zou, Le Yu, Yifei Huang +2

In many real-world scenarios (e.g., academic networks, social platforms), different types of entities are not only associated with texts but also connected by various relationships…

cs.CV2023

Memory-and-Anticipation Transformer for Online Action Understanding

Jiahao Wang, Guo Chen, Yifei Huang +2

Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal…

cs.CV202315 cited

VideoLLM: Modeling Video Sequence with Large Language Models

Guo Chen, Yin-Dong Zheng, Jiahao Wang +8

With the exponential growth of video data, there is an urgent need for automated technology to analyze and comprehend video content. However, existing video understanding models ar…

cs.CV20213 cited

Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips

Lijin Yang, Yifei Huang, Yusuke Sugano +1

First-person action recognition is a challenging task in video understanding. Because of strong ego-motion and a limited field of view, many backgrounds or noisy frames in a first-…

cs.CV20214 cited

Leveraging Human Selective Attention for Medical Image Analysis with Limited Training Data

Yifei Huang, Xiaoxiao Li, Lijin Yang +6

The human gaze is a cost-efficient physiological data that reveals human underlying attentional patterns. The selective attention mechanism helps the cognition system focus on task…