3 papers
cs.RO2026
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
Jialei Chen, Kai Wang, Kang Chen +9
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change th…
cs.CL2025
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
Zheyue Tan, Zhiyuan Li, Tao Yuan +13
Mixture-of-Experts (MoE) architectures have emerged as a promising approach to scale Large Language Models (LLMs). MoE boosts the efficiency by activating a subset of experts per t…
cs.LG2025
Megrez-Omni Technical Report
Boxun Li, Yadong Li, Zhiyuan Li +12
In this work, we present the Megrez models, comprising a language model (Megrez-3B-Instruct) and a multimodal model (Megrez-3B-Omni). These models are designed to deliver fast infe…