activity
20122025
most citedHET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework

57 citations · 171 across the 22 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.DC2025

TridentServe: A Stage-level Serving System for Diffusion Pipelines

Yifei Xia, Fangcheng Fu, Hao Yuan +6

Diffusion pipelines, renowned for their powerful visual generation capabilities, have seen widespread adoption in generative vision tasks (e.g., text-to-image/video). These pipelin…

cs.DC20252 cited

LobRA: Multi-tenant Fine-tuning over Heterogeneous Data

Sheng Lin, Fangcheng Fu, Haoyang Li +5

With the breakthrough of Transformer-based pre-trained models, the demand for fine-tuning (FT) to adapt the base pre-trained models to downstream applications continues to grow, so…

cs.LG20251 cited

Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-training

Minghao Xu, Jiaze Song, Keming Wu +3

Understanding the various properties of glycans with machine learning has shown some preliminary promise. However, previous methods mainly focused on modeling the backbone structur…

cs.LG2025

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling

Xiaodong Ji, Hailin Zhang, Fangcheng Fu +1

Many advanced Large Language Model (LLM) applications require long-context processing, but the self-attention module becomes a bottleneck during the prefilling stage of inference d…

cs.LG2025

Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately

Yuhang Wang, Youhe Jiang, Bin Cui +1

Recent advances in test-time scaling suggest that Large Language Models (LLMs) can gain better capabilities by generating Chain-of-Thought reasoning (analogous to human thinking) t…