6 papers
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
Zhiyuan Li, Linyuan Gao, Xuechun Ding +3
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem.…
AGGC: Adaptive Group Gradient Clipping for Stabilizing Large Language Model Training
Zhiyuan Li, Yuan Wu, Yi Chang
To stabilize the training of Large Language Models (LLMs), gradient clipping is a nearly ubiquitous heuristic used to alleviate exploding gradients. However, traditional global nor…
A Survey of Retentive Network
Haiqi Yang, Zhiyuan Li, Yi Chang +1
Retentive Network (RetNet) represents a significant advancement in neural network architecture, offering an efficient alternative to the Transformer. While Transformers rely on sel…
THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
Zhiyuan Li, Yi Chang, Yuan Wu
Large reasoning models (LRMs) have achieved impressive performance in complex tasks, often outperforming conventional large language models (LLMs). However, the prevalent issue of…
SegRet: An Efficient Design for Semantic Segmentation with Retentive Network
Zhiyuan Li, Yi Chang, Yuan Wu
With the rapid evolution of autonomous driving technology and intelligent transportation systems, semantic segmentation has become increasingly critical. Precise interpretation and…
A Survey of RWKV
Zhiyuan Li, Tingyu Xia, Yi Chang +1
The Receptance Weighted Key Value (RWKV) model offers a novel alternative to the Transformer architecture, merging the benefits of recurrent and attention-based systems. Unlike con…