4 papers
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
Jialei Chen, Kai Wang, Kang Chen +9
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change th…
Split Matching for Inductive Zero-shot Semantic Segmentation
Jialei Chen, Xu Zheng, Dongyue Li +6
Zero-shot Semantic Segmentation (ZSS) aims to segment categories that are not annotated during training. While fine-tuning vision-language models has achieved promising results, th…
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
Jialei Chen, Xu Zheng, Danda Pani Paudel +3
Zero-shot Semantic Segmentation (ZSS) aims to segment both seen and unseen classes using supervision from only seen classes. Beyond adaptation-based methods, distillation-based app…
BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
Jialei Chen, Xu Zheng, Danda Pani Paudel +3
Utilizing multi-modal data enhances scene understanding by providing complementary semantic and geometric information. Existing methods fuse features or distill knowledge from mult…