collaborators

5 papers

cs.RO2026

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

Jialei Chen, Kai Wang, Kang Chen +9

Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change th…

cs.CL2025

LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K

Tao Yuan, Xuefei Ning, Dong Zhou +10

State-of-the-art large language models (LLMs) are now claiming remarkable supported context lengths of 256k or even more. In contrast, the average context lengths of mainstream ben…

cs.CL2025

ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts

Zheyue Tan, Zhiyuan Li, Tao Yuan +13

Mixture-of-Experts (MoE) architectures have emerged as a promising approach to scale Large Language Models (LLMs). MoE boosts the efficiency by activating a subset of experts per t…

cs.CL2025

Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving

Yuxuan Zhou, Xien Liu, Chenwei Yan +8

Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored.…

cs.LG2025

Megrez-Omni Technical Report

Boxun Li, Yadong Li, Zhiyuan Li +12

In this work, we present the Megrez models, comprising a language model (Megrez-3B-Instruct) and a multimodal model (Megrez-3B-Omni). These models are designed to deliver fast infe…