activity
20192026
most citedRSPNet: Relative Speed Perception for Unsupervised Video Representation Learning

16 citations · 54 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

Towards Stable Cross-Domain Depression Recognition under Missing Modalities

Jiuyi Chen, Mingkui Tan, Haifeng Lu +4

Depression poses serious public health risks, including suicide, underscoring the urgency of timely and scalable screening. Multimodal automatic depression detection (ADD) offers a…

cs.CV2024

Towards Long Video Understanding via Fine-detailed Video Story Generation

Zeng You, Zhiquan Wen, Yaofo Chen +4

Long video understanding has become a critical task in computer vision, driving advancements across numerous applications from surveillance to content retrieval. Existing video und…

cs.CV20231 cited

Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Peihao Chen, Xinyu Sun, Hongyan Zhi +5

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by langu…

cs.CV202216 cited

Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation

Peihao Chen, Dongyu Ji, Kunyang Lin +4

We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions oft…

cs.CV202016 cited

RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning

Peihao Chen, Deng Huang, Dongliang He +5

We study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such…

cs.CV20204 cited

Location-aware Graph Convolutional Networks for Video Question Answering

Deng Huang, Peihao Chen, Runhao Zeng +3

We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art method…