collaborators

9 papers

cs.CL2026

HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents

Jiangze Yan, Yi Shen, Wenjing Zhang +5

Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be re…

cs.CL2026

Mixture of Heterogeneous Grouped Experts for Language Modeling

Zhicheng Ma, Xiang Liu, Zhaoxiang Liu +5

Large Language Models (LLMs) based on Mixture-of-Experts (MoE) are pivotal in industrial applications for their ability to scale performance efficiently. However, standard MoEs enf…

cs.DC2026

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

Minjie Hua, Ning Wang, Peijun Yang +2

OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window. For this worklo…

cs.AI2026

HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation

Wenjing Zhang, Jiangze Yan, Jieyun Huang +7

Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitation of rejection sampling. Standard methods treat th…

cs.LG2026

DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models

Yi Shen, Jian Zhang, Jieyun Huang +7

Recent advancements in slow thinking reasoning models have shown exceptional performance in complex reasoning tasks. However, these models often exhibit overthinking (generating re…

cs.LG2025

Quantitative Analysis of Performance Drop in DeepSeek Model Quantization

Enbo Zhao, Yi Shen, Shuming Shi +7

Recently, there is a high demand for deploying DeepSeek-R1 and V3 locally, possibly because the official service often suffers from being busy and some organizations have data priv…