papers

Publications (35)

cs.CR2025

PriFFT: Privacy-preserving Federated Fine-tuning of Large Language Models via Hybrid Secret Sharing

Zhichao You, Xuewen Dong, Ke Cheng +5

Fine-tuning large language models (LLMs) raises privacy concerns due to the risk of exposing sensitive training data. Federated learning (FL) mitigates this risk by keeping trainin…

physics.optics2025

Error-Corrected Eternal Lifetime Storage

Jie Ma, Chu-Han Wang, Xiao-Yun Xu +5

In the information explosion era, the demand for high-density stable storage technologies is soaring. Multi-dimensional optical storage with femtosecond laser writing offers a pote…

cs.DC2025

P/D-Device: Disaggregated Large Language Model between Cloud and Devices

Yibo Jin, Yixu Xu, Yue Chen +27

Serving disaggregated large language models has been widely adopted in industrial practice for enhanced performance. However, too many tokens generated in decoding phase, i.e., occ…

cs.DC2025

DecLock: A Case of Decoupled Locking for Disaggregated Memory

Hanze Zhang, Ke Cheng, Rong Chen +2

This paper reveals that locking can significantly degrade the performance of applications on disaggregated memory (DM), sometimes by several orders of magnitude, due to contention…

cs.CR2026

FedGSA: Geometry-Consistent Subspace Aggregation for Differentially Private Federated LoRA

Lele Zheng, Ruijie Hu, Tao Zhang +2

Low-Rank Adaptation (LoRA) enables communication-efficient federated fine-tuning of pretrained language models. However, integrating differential privacy (DP) into federated LoRA r…

cond-mat.mtrl-sci2011

Strong Visible Absorption and Photoluminescence of Titanic Acid Nanotubes by Hydrothermal Method

Baoli Tian, Xing Zhang, Shuxi Dai +6

Titanic acid nanotubes (with a chemical formula H2Ti2O4(OH)2, abbreviated as TANTs) were synthesized by the hydrothermal method using commercial TiO2 nanoparticle powder (P25, Degu…

cs.CV2019

Patch Transformer for Multi-tagging Whole Slide Histopathology Images

Weijian Li, Viet-Duy Nguyen, Haofu Liao +3

Automated whole slide image (WSI) tagging has become a growing demand due to the increasing volume and diversity of WSIs collected nowadays in histopathology. Various methods have…

cs.CR2026

When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents

Di Lu, Yongzhi Liao, Xutong Mu +5

Host-acting agents promise a convenient interaction model in which users specify goals and the system determines how to realize them. We argue that this convenience introduces a di…

cs.LG2026

Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models

Lele Zheng, Weifeng Kong, Xinyi Zhang +3

Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-…

cs.LG2026

PRISM: Parallel Residual Iterative Sequence Model

Jie Jiang, Ke Cheng, Xin Xu +8

Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models. Existing efficient architectures are…

cs.LG2024

Co-Neighbor Encoding Schema: A Light-cost Structure Encoding Method for Dynamic Link Prediction

Ke Cheng, Linzhi Peng, Junchen Ye +2

Structure encoding has proven to be the key feature to distinguishing links in a graph. However, Structure encoding in the temporal graph keeps changing as the graph evolves, repea…

cs.LG2023

SMAP: A Novel Heterogeneous Information Framework for Scenario-based Optimal Model Assignment

Zekun Qiu, Zhipu Xie, Zehua Ji +2

The increasing maturity of big data applications has led to a proliferation of models targeting the same objectives within the same scenarios and datasets. However, selecting the m…

cs.CV2020

Compact Global Descriptor for Neural Networks

Xiangyu He, Ke Cheng, Qiang Chen +3

Long-range dependencies modeling, widely used in capturing spatiotemporal correlation, has shown to be effective in CNN dominated computer vision tasks. Yet neither stacks of convo…

cs.LG2025

MSCMNet: Multi-scale Semantic Correlation Mining for Visible-Infrared Person Re-Identification

Xuecheng Hua, Ke Cheng, Hu Lu +3

The main challenge in the Visible-Infrared Person Re-Identification (VI-ReID) task lies in how to extract discriminative features from different modalities for matching purposes. W…

cs.LG2024

DyGKT: Dynamic Graph Learning for Knowledge Tracing

Ke Cheng, Linzhi Peng, Pengyang Wang +3

Knowledge Tracing aims to assess student learning states by predicting their performance in answering questions. Different from the existing research which utilizes fixed-length le…

physics.ins-det2014

Application of the DRS4 Chip for GHz Waveform Digitizing Circuit

HaiBo Yang, Hong Su, Jie Kong +4

At present, fast waveform digitizing circuit is more and more employed in modern physics experiments for processing the signals from an array detector. A new fast waveform sampling…

cs.CL2025

CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences

Ziran Qin, Yuchen Cao, Mingbao Lin +5

Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference bu…

cs.DC2024

Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction

Ke Cheng, Wen Hu, Zhi Wang +3

Nowadays, large language models (LLMs) are published as a service and can be accessed by various applications via APIs, also known as language-model-as-a-service (LMaaS). Without k…

cs.CV2024

DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model

Yuqi Wang, Ke Cheng, Jiawei He +5

Driving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashe…

cs.DC2025

FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference

Bingzhe Zhao, Ke Cheng, Aomufei Yuan +5

KV cache techniques in Transformer models aim to reduce redundant computations at the expense of substantially increased memory usage, making KV cache compression an important and…

cs.IR2026

Recurrent Preference Memory for Efficient Long-Sequence Generative Recommendation

Yixiao Chen, Yuan Wang, Yue Liu +9

Generative recommendation (GenRec) models typically model user behavior via full attention, but scaling to lifelong sequences is hindered by prohibitive computational costs and noi…

cs.CR2025

Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset

Anxiao Song, Shujie Cui, Jianli Bai +3

In light of increasing privacy concerns and stringent legal regulations, using secure multiparty computation (MPC) to enable collaborative GBDT model training among multiple data o…

physics.optics2024

Ultranarrow-linewidth Wavelength-Vortex Metasurface Holography

Weijia Meng, Johannes E. Fröch, Ke Cheng +6

Ultrathin metasurface holograms, with thicknesses comparable to the operating wavelength, leverage multiple degrees of freedom of light to address independent image channels, there…

cs.LG2021

FedProc: Prototypical Contrastive Federated Learning on Non-IID data

Xutong Mu, Yulong Shen, Ke Cheng +4

Federated learning allows multiple clients to collaborate to train high-performance deep learning models while keeping the training data locally. However, when the local data of al…

physics.ins-det2015

Analysis of digital timing methods with DRS4 module

Cheng-Ming Du, Jin-Da Chen, Xiu-Ling Zhang +7

A new Digital Pulse Processing (DPP) module has been developed, based on a domino ring sampler version 4 chip (DRS4), with good time resolution for LaBr3 detectors, and different d…

physics.app-ph2024

Symmetry engineering in 2D bioelectronics facilitating augmented biosensing interfaces

Yizhang Wu, Yihan Liu, Yuan Li +13

Symmetry lies at the heart of 2D bioelectronics, determining material properties at the fundamental level. Breaking the symmetry allows emergent functionalities and effects. Howeve…

cs.CV2026

VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning

Li-Heng Chen, Ke Cheng, Yahui Liu +3

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…

cs.DC2025

Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Ke Cheng, Wen Hu, Zhi Wang +3

Large language models (LLMs) iteratively generate text token by token, with memory usage increasing with the length of generated token sequences. Since the request generation lengt…

cs.LG2026

Differentially Private Subspace Fine-Tuning for Large Language Models

Lele Zheng, Xiang Wang, Tao Zhang +3

Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differenti…

cs.CV2022

PKD: General Distillation Framework for Object Detectors via Pearson Correlation Coefficient

Weihan Cao, Yifan Zhang, Jianfei Gao +3

Knowledge distillation(KD) is a widely-used technique to train compact models in object detection. However, there is still a lack of study on how to distill between heterogeneous d…

cs.CV2026

MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

Ke Cheng, Hanqiao Ye, Lei Shi +8

The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…

#driving simulation#multimodal generation#object manipulation#long-horizon streaming
cs.IR2026

HyTRec: A Hybrid Temporal-Aware Attention Architecture for Long Behavior Sequential Recommendation

Lei Xin, Yuhao Zheng, Ke Cheng +3

Modeling long sequences of user behaviors has emerged as a critical frontier in generative recommendation. However, existing solutions face a dilemma: linear attention mechanisms a…

cs.DC2025

SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines

Ke Cheng, Zhi Wang, Wen Hu +3

As large language models (LLMs) are gaining increasing popularity across a wide range of web applications, it is of great importance to optimize service-level objectives (SLOs) for…

physics.ins-det2015

Study of time resolution by digital methods with DRS4 system

Cheng-Ming Du, Jin-Da Chen, Xiu-Ling Zhang +5

A new Digital Pulse Processing (DPP) system, based on domino ring sampler version 4 (DRS4), with good time resolution for the LaBr3 detectors has been developed and different digit…

cs.DC2026

RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference

Jiarui Wang, Huichao Chai, Yuanhang Zhang +38

Real-time recommender systems execute multi-stage cascades (retrieval, pre-processing, fine-grained ranking) under strict tail-latency SLOs, leaving only tens of milliseconds for r…