4 papers
LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment
Huaihai Lyu, Chaofan Chen, Yuheng Ji +4
We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action representations compatible with the s…
DMC: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
Jiayi Zou, Chaofan Chen, Bing-Kun Bao +1
Egocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos…
OmniSAT: Compact Action Token, Faster Auto Regression
Huaihai Lyu, Chaofan Chen, Senwei Xie +4
Existing Vision-Language-Action (VLA) models can be broadly categorized into diffusion-based and auto-regressive (AR) approaches: diffusion models capture continuous action distrib…
EgoPrompt: Prompt Learning for Egocentric Action Recognition
Huaihai Lyu, Chaofan Chen, Yuheng Ji +1
Driven by the increasing demand for applications in augmented and virtual reality, egocentric action recognition has emerged as a prominent research area. It is typically divided i…