Publications (8)
Predicting 30-Day Hospital Readmission in Medicare Patients: Insights from an LSTM Deep Learning Model
Xintao Li, Sibei Liu, Dezhi Yu +2
Readmissions among Medicare beneficiaries are a major problem for the US healthcare system from a perspective of both healthcare operations and patient caregiving outcomes. Our stu…
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
Yang Wu, Huayi Zhang, Yizheng Jiao +6
Instruction tuning has underscored the significant potential of large language models (LLMs) in producing more human controllable and effective outputs in various domains. In this…
DK-RRT: Deep Koopman RRT for Collision-Aware Motion Planning of Space Manipulators in Dynamic Debris Environments
Qi Chen, Rui Liu, Kangtong Mo +2
Trajectory planning for robotic manipulators operating in dynamic orbital debris environments poses significant challenges due to complex obstacle movements and uncertainties. This…
Direct Preference Optimization for Chatbot Fine-Tuning: An Empirical Study
Dezhi Yu, Yvonne Qiu, ShuoJia Fu
We present an approach to fine-tuning large language models using Direct Preference Optimization (DPO), a reinforcement learning technique. Our experimental results demonstrate tha…
Research on reinforcement learning based warehouse robot navigation algorithm in complex warehouse layout
Keqin Li, Lipeng Liu, Jiajing Chen +5
In this paper, how to efficiently find the optimal path in complex warehouse layout and make real-time decision is a key problem. This paper proposes a new method of Proximal Polic…
AdaMixup: A Dynamic Defense Framework for Membership Inference Attack Mitigation
Ying Chen, Jiajing Chen, Yijie Weng +3
Membership inference attacks have emerged as a significant privacy concern in the training of deep learning models, where attackers can infer whether a data point was part of the t…
KVDirect: Distributed Disaggregated LLM Inference
Shiyang Chen, Rain Jiang, Dezhi Yu +6
Large Language Models (LLMs) have become the new foundation for many applications, reshaping human society like a storm. Disaggregated inference, which separates prefill and decode…
Generative Modeling of Bach-Style Symbolic Music: A Comparative Study of Autoregressive, Latent-Variable, and Adversarial Approaches
Dezhi Yu, Kyuil Lee, Yongkang Huang
We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models inc…