3 papers
cs.CV2025
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
Jinghan Zhao, Yifei Huang, Feng Lu
Learning procedural-aware video representations is a key step towards building agents that can reason about and execute complex tasks. Existing methods typically address this probl…
cs.CV2025
Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
Zefeng Qian, Xincheng Yao, Yifei Huang +3
Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recen…
cs.CV2025
Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition
Zefeng Qian, Chongyang Zhang, Yifei Huang +2
Few-shot Action Recognition (FSAR) constitutes a crucial challenge in computer vision, entailing the recognition of actions from a limited set of examples. Recent approaches mainly…