3 papers
cs.LG2026
Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies
Wei Jiang, Wei Wang
Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Existing efficiency methods mainly reduce vi…
cs.CV2026
Temporal Prototyping and Hierarchical Alignment for Unsupervised Video-based Visible-Infrared Person Re-Identification
Zhiyong Li, Wei Jiang, Haojie Liu +3
Visible-infrared person re-identification (VI-ReID) enables cross-modality identity matching for all-day surveillance, yet existing methods predominantly focus on the image level o…
cs.CL2025
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
Pu Zhao, Arash Akbari, Xuan Shen +16
Recently, Large Language Models (LLMs) have undergone a significant transformation, marked by a rapid rise in both their popularity and capabilities. Leading this evolution are pro…