Publications (35)
PriFFT: Privacy-preserving Federated Fine-tuning of Large Language Models via Hybrid Secret Sharing
Zhichao You, Xuewen Dong, Ke Cheng +5
Fine-tuning large language models (LLMs) raises privacy concerns due to the risk of exposing sensitive training data. Federated learning (FL) mitigates this risk by keeping trainin…
Error-Corrected Eternal Lifetime Storage
Jie Ma, Chu-Han Wang, Xiao-Yun Xu +5
In the information explosion era, the demand for high-density stable storage technologies is soaring. Multi-dimensional optical storage with femtosecond laser writing offers a pote…
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
Yibo Jin, Yixu Xu, Yue Chen +27
Serving disaggregated large language models has been widely adopted in industrial practice for enhanced performance. However, too many tokens generated in decoding phase, i.e., occ…
DecLock: A Case of Decoupled Locking for Disaggregated Memory
Hanze Zhang, Ke Cheng, Rong Chen +2
This paper reveals that locking can significantly degrade the performance of applications on disaggregated memory (DM), sometimes by several orders of magnitude, due to contention…
FedGSA: Geometry-Consistent Subspace Aggregation for Differentially Private Federated LoRA
Lele Zheng, Ruijie Hu, Tao Zhang +2
Low-Rank Adaptation (LoRA) enables communication-efficient federated fine-tuning of pretrained language models. However, integrating differential privacy (DP) into federated LoRA r…
Strong Visible Absorption and Photoluminescence of Titanic Acid Nanotubes by Hydrothermal Method
Baoli Tian, Xing Zhang, Shuxi Dai +6
Titanic acid nanotubes (with a chemical formula H2Ti2O4(OH)2, abbreviated as TANTs) were synthesized by the hydrothermal method using commercial TiO2 nanoparticle powder (P25, Degu…
Patch Transformer for Multi-tagging Whole Slide Histopathology Images
Weijian Li, Viet-Duy Nguyen, Haofu Liao +3
Automated whole slide image (WSI) tagging has become a growing demand due to the increasing volume and diversity of WSIs collected nowadays in histopathology. Various methods have…
When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents
Di Lu, Yongzhi Liao, Xutong Mu +5
Host-acting agents promise a convenient interaction model in which users specify goals and the system determines how to realize them. We argue that this convenience introduces a di…
Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models
Lele Zheng, Weifeng Kong, Xinyi Zhang +3
Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-…
PRISM: Parallel Residual Iterative Sequence Model
Jie Jiang, Ke Cheng, Xin Xu +8
Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models. Existing efficient architectures are…
Co-Neighbor Encoding Schema: A Light-cost Structure Encoding Method for Dynamic Link Prediction
Ke Cheng, Linzhi Peng, Junchen Ye +2
Structure encoding has proven to be the key feature to distinguishing links in a graph. However, Structure encoding in the temporal graph keeps changing as the graph evolves, repea…
SMAP: A Novel Heterogeneous Information Framework for Scenario-based Optimal Model Assignment
Zekun Qiu, Zhipu Xie, Zehua Ji +2
The increasing maturity of big data applications has led to a proliferation of models targeting the same objectives within the same scenarios and datasets. However, selecting the m…
Compact Global Descriptor for Neural Networks
Xiangyu He, Ke Cheng, Qiang Chen +3
Long-range dependencies modeling, widely used in capturing spatiotemporal correlation, has shown to be effective in CNN dominated computer vision tasks. Yet neither stacks of convo…
MSCMNet: Multi-scale Semantic Correlation Mining for Visible-Infrared Person Re-Identification
Xuecheng Hua, Ke Cheng, Hu Lu +3
The main challenge in the Visible-Infrared Person Re-Identification (VI-ReID) task lies in how to extract discriminative features from different modalities for matching purposes. W…
DyGKT: Dynamic Graph Learning for Knowledge Tracing
Ke Cheng, Linzhi Peng, Pengyang Wang +3
Knowledge Tracing aims to assess student learning states by predicting their performance in answering questions. Different from the existing research which utilizes fixed-length le…
Application of the DRS4 Chip for GHz Waveform Digitizing Circuit
HaiBo Yang, Hong Su, Jie Kong +4
At present, fast waveform digitizing circuit is more and more employed in modern physics experiments for processing the signals from an array detector. A new fast waveform sampling…
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
Ziran Qin, Yuchen Cao, Mingbao Lin +5
Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference bu…
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
Ke Cheng, Wen Hu, Zhi Wang +3
Nowadays, large language models (LLMs) are published as a service and can be accessed by various applications via APIs, also known as language-model-as-a-service (LMaaS). Without k…
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
Yuqi Wang, Ke Cheng, Jiawei He +5
Driving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashe…
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
Bingzhe Zhao, Ke Cheng, Aomufei Yuan +5
KV cache techniques in Transformer models aim to reduce redundant computations at the expense of substantially increased memory usage, making KV cache compression an important and…
Recurrent Preference Memory for Efficient Long-Sequence Generative Recommendation
Yixiao Chen, Yuan Wang, Yue Liu +9
Generative recommendation (GenRec) models typically model user behavior via full attention, but scaling to lifelong sequences is hindered by prohibitive computational costs and noi…
Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset
Anxiao Song, Shujie Cui, Jianli Bai +3
In light of increasing privacy concerns and stringent legal regulations, using secure multiparty computation (MPC) to enable collaborative GBDT model training among multiple data o…
Ultranarrow-linewidth Wavelength-Vortex Metasurface Holography
Weijia Meng, Johannes E. Fröch, Ke Cheng +6
Ultrathin metasurface holograms, with thicknesses comparable to the operating wavelength, leverage multiple degrees of freedom of light to address independent image channels, there…
FedProc: Prototypical Contrastive Federated Learning on Non-IID data
Xutong Mu, Yulong Shen, Ke Cheng +4
Federated learning allows multiple clients to collaborate to train high-performance deep learning models while keeping the training data locally. However, when the local data of al…
Analysis of digital timing methods with DRS4 module
Cheng-Ming Du, Jin-Da Chen, Xiu-Ling Zhang +7
A new Digital Pulse Processing (DPP) module has been developed, based on a domino ring sampler version 4 chip (DRS4), with good time resolution for LaBr3 detectors, and different d…
Symmetry engineering in 2D bioelectronics facilitating augmented biosensing interfaces
Yizhang Wu, Yihan Liu, Yuan Li +13
Symmetry lies at the heart of 2D bioelectronics, determining material properties at the fundamental level. Breaking the symmetry allows emergent functionalities and effects. Howeve…
VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning
Li-Heng Chen, Ke Cheng, Yahui Liu +3
Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
Ke Cheng, Wen Hu, Zhi Wang +3
Large language models (LLMs) iteratively generate text token by token, with memory usage increasing with the length of generated token sequences. Since the request generation lengt…
Differentially Private Subspace Fine-Tuning for Large Language Models
Lele Zheng, Xiang Wang, Tao Zhang +3
Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differenti…
PKD: General Distillation Framework for Object Detectors via Pearson Correlation Coefficient
Weihan Cao, Yifan Zhang, Jianfei Gao +3
Knowledge distillation(KD) is a widely-used technique to train compact models in object detection. However, there is still a lack of study on how to distill between heterogeneous d…
MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
Ke Cheng, Hanqiao Ye, Lei Shi +8
The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…
HyTRec: A Hybrid Temporal-Aware Attention Architecture for Long Behavior Sequential Recommendation
Lei Xin, Yuhao Zheng, Ke Cheng +3
Modeling long sequences of user behaviors has emerged as a critical frontier in generative recommendation. However, existing solutions face a dilemma: linear attention mechanisms a…
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
Ke Cheng, Zhi Wang, Wen Hu +3
As large language models (LLMs) are gaining increasing popularity across a wide range of web applications, it is of great importance to optimize service-level objectives (SLOs) for…
Study of time resolution by digital methods with DRS4 system
Cheng-Ming Du, Jin-Da Chen, Xiu-Ling Zhang +5
A new Digital Pulse Processing (DPP) system, based on domino ring sampler version 4 (DRS4), with good time resolution for the LaBr3 detectors has been developed and different digit…
RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference
Jiarui Wang, Huichao Chai, Yuanhang Zhang +38
Real-time recommender systems execute multi-stage cascades (retrieval, pre-processing, fine-grained ranking) under strict tail-latency SLOs, leaving only tens of milliseconds for r…