activity
20242026
collaborators

14 papers

cs.AI2026

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Lei Xin, Bin Gu, Peize Li +9

Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. W…

cs.RO2026

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

Jun Liu, Pu Zhao, Zhenglun Kong +12

Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the en…

cs.LG2026

Structured Agent Distillation for Large Language Model

Jun Liu, Zhenglun Kong, Peiyan Dong +10

Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-style frameworks. Yet, their practical de…

cs.CL2026

Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment

Jun Liu, Zhenglun Kong, Pu Zhao +9

Structured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redunda…

cs.CL2026

Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA

Pu Zhao, Arash Akbari, Xuan Shen +16

Recently, Large Language Models (LLMs) have undergone a significant transformation, marked by a rapid rise in both their popularity and capabilities. Leading this evolution are pro…

cs.CV2025

TSLA: A Task-Specific Learning Adaptation for Semantic Segmentation on Autonomous Vehicles Platform

Jun Liu, Zhenglun Kong, Pu Zhao +9

Autonomous driving platforms encounter diverse driving scenarios, each with varying hardware resources and precision requirements. Given the computational limitations of embedded d…