4 papers
TADS: Task-Aware Data Selection for Multi-Task Multimodal Pre-Training
Guanjie Cheng, Boyi Li, Lingyu Sun +4
Large-scale multimodal pre-trained models like CLIP rely heavily on high-quality training data, yet raw web-crawled datasets are often noisy, misaligned, and redundant, leading to…
TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain
Yidan Sun, Mengying Zhu, Feiyue Chen +5
Large language models (LLMs) have demonstrated impressive performance in text generation tasks; however, their embedding spaces often suffer from the isotropy problem, resulting in…
DP-GENG : Differentially Private Dataset Distillation Guided by DP-Generated Data
Shuo Shi, Jinghuai Zhang, Shijie Jiang +5
Dataset distillation (DD) compresses large datasets into smaller ones while preserving the performance of models trained on them. Although DD is often assumed to enhance data priva…
LSHFed: Robust and Communication-Efficient Federated Learning with Locally-Sensitive Hashing Gradient Mapping
Guanjie Cheng, Mengzhen Yang, Xinkui Zhao +5
Federated learning (FL) enables collaborative model training across distributed nodes without exposing raw data, but its decentralized nature makes it vulnerable in trust-deficient…