collaborators

9 papers

cs.RO2026

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de…

cs.CV2026

Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition

Hongyu Qu, Ling Xing, Jiachao Zhang +3

Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level representations for each video by desi…

cs.CV2026

See the Text: From Tokenization to Visual Reading

Ling Xing, Rui Yan, Alex Jinpeng Wang +2

People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle ty…

cs.CV2026

Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition

Hongyu Qu, Xiangbo Shu, Rui Yan +3

Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coar…

cs.CV2025

OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild

Hongyu Qu, Jianan Wei, Xiangbo Shu +3

Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to i) the scarcity of annotated datasets, and ii) the insufficient diversity of…

cs.CV2025

Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization

Ling Xing, Hongyu Qu, Rui Yan +2

Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where even…