14 papers
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Yutong He, Daibo Li, Guohong Li +15
ReasFlow is an autonomous multi‑agent system that leverages large language models to perform rigorous mathematical reasoning, retrieve relevant knowledge, and generate complete res…
CentroidKV: Efficient Long-Context LLM Inference via KV Cache Clustering
Jie Hu, Shengnan Wang, Yutong He +8
Large language models (LLMs) with extended context windows have become increasingly prevalent for tackling complex tasks. However, the substantial Key-Value (KV) cache required for…
FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection
Yutong He, Zhengyang Huang, Jiahe Geng +1
Federated learning enables a population of clients to collaboratively train machine learning models without exchanging their raw data, but standard algorithms such as FedAvg suffer…
DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation
Jie Hu, Zixiang Gao, Yutong He +1
Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational co…
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
Yuxi Liu, Zekun Zhang, Yixiang Cai +3
Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, yet their attention complexity poses a formidable bottleneck for long-sequence…
TAH-QUANT: Effective Activation Quantization in Pipeline Parallelism over Slow Network
Guangxin He, Yuan Cao, Yutong He +4
Decentralized training of large language models offers the opportunity to pool computational resources across geographically distributed participants, but is often bottlenecked by…