1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2026
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
Gradwell Dzikanyanga, Weihao Yang, Hao Huang +4
Key-value (KV) caching is critical for efficient inference in large language models (LLMs), yet its memory footprint scales linearly with context length, resulting in a severe scal…
cs.AI2026
Memory as Asset: From Agent-centric to Human-centric Memory Management
Yanqi Pan, Qinghao Huang, Weihao Yang
We proudly introduce Memory-as-Asset, a new memory paradigm towards human-centric artificial general intelligence (AGI). In this paper, we formally emphasize that human-centric, pe…
cs.DC2025★ 1 cited
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
Weihao Yang, Hao Huang, Donglei Wu +6
Mixture-of-Experts (MoE) has become a popular architecture for scaling large models. However, the rapidly growing scale outpaces model training on a single DC, driving a shift towa…