activity
20242026
collaborators

9 papers

cs.CV2026

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

Silpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie +2

Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses thi…

cs.CL2026

Measuring Stereotype and Deviation Biases in Large Language Models

Daniel Wang, Eli Brignac, Minjia Mao +1

Large language models (LLMs) are widely applied across diverse domains, raising concerns about their limitations and potential risks. In this study, we investigate two types of bia…

cs.LG2025

Synthetic Series-Symbol Data Generation for Time Series Foundation Models

Wenxuan Wang, Kai Wu, Yujian Betterest Li +2

Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as training data scarcity and imbalance continue to hinder their dev…

cs.CL2025

KG2QA: Knowledge Graph-enhanced Retrieval-augmented Generation for Communication Standards Question Answering

Zhongze Luo, Weixuan Wan, Tianya Zhang +2

The rapid evolution of communication technologies has led to an explosion of standards, rendering traditional expert-dependent consultation methods inefficient and slow. To address…

cs.CL2025

Fine Tuning Methods for Low-resource Languages

Tim Bakkenes, Daniel Wang, Anton Johansson

The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trained on English texts and culture which makes them underperform in other language…

cs.LG2025

Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation

Brian Shing-Hei Wong, Joshua Mincheol Kim, Sin-Hang Fung +11

Allergens, typically proteins capable of triggering adverse immune responses, represent a significant public health challenge. To accurately identify allergen proteins, we introduc…