collaborators

10 papers

cs.RO2026

LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models

Zhengyan Qian, Rui Yan, Alex Jinpeng Wang +1

Visual cues are increasingly adopted to guide robot learning, but whether Vision-Language-Action (VLA) models can reliably follow authorized cues while disregarding unauthorized on…

cs.RO2026

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de…

cs.CV2026

EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond

Meiqi Cao, Xiangbo Shu, Jiachao Zhang +3

Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current l…

cs.CL2026

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Liujie Zhang, Benzhe Ning, Rui Yang +8

Reinforcement learning (RL) post-training has proven effective at unlocking reasoning, self-reflection, and tool-use capabilities in large language models. As models extend to omni…

cs.CV2026

See the Text: From Tokenization to Visual Reading

Ling Xing, Rui Yan, Alex Jinpeng Wang +2

People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle ty…

cs.CV2026

Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition

Hongyu Qu, Xiangbo Shu, Rui Yan +3

Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coar…