collaborators

6 papers

cs.AI2026

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Zhiyuan Li, Linyuan Gao, Xuechun Ding +3

Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem.…

cs.LG2026

AGGC: Adaptive Group Gradient Clipping for Stabilizing Large Language Model Training

Zhiyuan Li, Yuan Wu, Yi Chang

To stabilize the training of Large Language Models (LLMs), gradient clipping is a nearly ubiquitous heuristic used to alleviate exploding gradients. However, traditional global nor…

cs.CL2025

A Survey of Retentive Network

Haiqi Yang, Zhiyuan Li, Yi Chang +1

Retentive Network (RetNet) represents a significant advancement in neural network architecture, offering an efficient alternative to the Transformer. While Transformers rely on sel…

cs.CL2025

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Zhiyuan Li, Yi Chang, Yuan Wu

Large reasoning models (LRMs) have achieved impressive performance in complex tasks, often outperforming conventional large language models (LLMs). However, the prevalent issue of…

eess.IV2025

SegRet: An Efficient Design for Semantic Segmentation with Retentive Network

Zhiyuan Li, Yi Chang, Yuan Wu

With the rapid evolution of autonomous driving technology and intelligent transportation systems, semantic segmentation has become increasingly critical. Precise interpretation and…

cs.CL2025

A Survey of RWKV

Zhiyuan Li, Tingyu Xia, Yi Chang +1

The Receptance Weighted Key Value (RWKV) model offers a novel alternative to the Transformer architecture, merging the benefits of recurrent and attention-based systems. Unlike con…