9 papers
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
Silpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie +2
Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses thi…
Measuring Stereotype and Deviation Biases in Large Language Models
Daniel Wang, Eli Brignac, Minjia Mao +1
Large language models (LLMs) are widely applied across diverse domains, raising concerns about their limitations and potential risks. In this study, we investigate two types of bia…
Synthetic Series-Symbol Data Generation for Time Series Foundation Models
Wenxuan Wang, Kai Wu, Yujian Betterest Li +2
Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as training data scarcity and imbalance continue to hinder their dev…
KG2QA: Knowledge Graph-enhanced Retrieval-augmented Generation for Communication Standards Question Answering
Zhongze Luo, Weixuan Wan, Tianya Zhang +2
The rapid evolution of communication technologies has led to an explosion of standards, rendering traditional expert-dependent consultation methods inefficient and slow. To address…
Fine Tuning Methods for Low-resource Languages
Tim Bakkenes, Daniel Wang, Anton Johansson
The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trained on English texts and culture which makes them underperform in other language…
Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation
Brian Shing-Hei Wong, Joshua Mincheol Kim, Sin-Hang Fung +11
Allergens, typically proteins capable of triggering adverse immune responses, represent a significant public health challenge. To accurately identify allergen proteins, we introduc…