9 papers
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de…
Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition
Hongyu Qu, Ling Xing, Jiachao Zhang +3
Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level representations for each video by desi…
See the Text: From Tokenization to Visual Reading
Ling Xing, Rui Yan, Alex Jinpeng Wang +2
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle ty…
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
Hongyu Qu, Xiangbo Shu, Rui Yan +3
Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coar…
OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
Hongyu Qu, Jianan Wei, Xiangbo Shu +3
Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to i) the scarcity of annotated datasets, and ii) the insufficient diversity of…
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
Ling Xing, Hongyu Qu, Rui Yan +2
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where even…