8 citations · 29 across the 16 of their papers we have counts for
19 papers
Convex Dataset Valuation for Post-Training
Siqi Zeng, Christopher Jung, Rui Li +7
Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers face constraints on compute, l…
Multimodal Generative Recommendation for Fusing Semantic and Collaborative Signals
Moritz Vandenhirtz, Kaveh Hassani, Shervin Ghasemlou +5
Sequential recommender systems rank relevant items by modeling a user's interaction history and computing the inner product between the resulting user representation and stored ite…
Towards measuring fairness in speech recognition: Fair-Speech dataset
Irina-Elena Veliche, Zhuangqun Huang, Vineeth Ayyat Kochaniyan +3
The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper…
Mitigating Unintended Memorization in Language Models via Alternating Teaching
Zhe Liu, Xuedong Zhang, Fuchun Peng
Recent research has shown that language models have a tendency to memorize rare or unique sequences in the training corpora which can thus leak sensitive attributes of user data. W…
Group Personalized Federated Learning
Zhe Liu, Yue Hui, Fuchun Peng
Federated learning (FL) can help promote data privacy by training a shared model in a de-centralized manner on the physical devices of clients. In the presence of highly heterogene…
Modeling Dependent Structure for Utterances in ASR Evaluation
Zhe Liu, Fuchun Peng
The bootstrap resampling method has been popular for performing significance analysis on word error rate (WER) in automatic speech recognition (ASR) evaluation. To deal with depend…