collaborators

9 papers

cs.IR2026

Memory Layer: Train the In-Model Cache for Recommendation Models

Liangyuan Na, Gufan Yin, Yixin Bao +19

Early ranking stages in recommendation systems precompute item embeddings and cache them in-model for scoring within strict latency constraints. Because this cache exists only at s…

cs.DC2026

Tropical: Enhancing SLO Attainment in Disaggregated LLM Serving via SLO-Aware Multiplexing

Jinming Ma, Jiefei Chen, Xiuhong Li +5

To guarantee service quality in transformer based large language model (LLM) serving, it is essential to meet the latency constraints of both the prefill phase (measured by Time-to…

cs.IR2026

SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs

Bi Xue, Hong Wu, Lei Chen +29

Serving deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from…

cs.DC2026

W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs

Yuanhong He, Peiyu Niu, Jun Chen +2

As Large Language Models (LLMs) scale, weight-only quantization (W4A16: 4-bit weights, 16-bit activations) becomes critical for reducing memory footprint with minimal accuracy loss…

cs.CV2026

Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers

Chaojie Yang, Tian Li, Yue Zhang +1

Diffusion Transformer (DiT) architectures have significantly advanced Text-to-Image (T2I) generation but suffer from prohibitive computational costs and deployment barriers. To add…

cs.SD2025

ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis

Stephen Ni-Hahn, Chao Péter Yang, Mingchen Ma +3

Artificial Intelligence (AI) for music generation is undergoing rapid developments, with recent symbolic models leveraging sophisticated deep learning and diffusion model algorithm…