5 papers
SafeImpute: Reliable Clinical Data Imputation via Conformal Selection
Xinrui He, Mengting Ai, Junting Wang +2
Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness. While many im…
Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
Tianxin Wei, Yifan Chen, Xinrui He +2
Distribution shifts between training and testing samples frequently occur in practice and impede model generalization performance. This crucial challenge thereby motivates studies…
RAG over Tables: Hierarchical Memory Index, Multi-Stage Retrieval, and Benchmarking
Jiaru Zou, Dongqi Fu, Sirui Chen +5
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating them with an external knowledge base to improve the answer relevance and accuracy. In real…
Dataset Distillation via the Wasserstein Metric
Haoyang Liu, Yijiang Li, Tiancheng Xing +5
Dataset Distillation (DD) aims to generate a compact synthetic dataset that enables models to achieve performance comparable to training on the full large dataset, significantly re…
Co-clustering for Federated Recommender System
Xinrui He, Shuo Liu, Jackey Keung +1
As data privacy and security attract increasing attention, Federated Recommender System (FRS) offers a solution that strikes a balance between providing high-quality recommendation…