6 papers
DeFrame: Debiasing Large Language Models Against Framing Effects
Kahee Lim, Soyeon Kim, Steven Euijong Whang
As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has become crucial. Despite many efforts, an…
Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
Soyeon Kim, Jindong Wang, Xing Xie +1
Facts change over time, making it essential for Large Language Models (LLMs) to handle time-sensitive factual knowledge accurately and reliably. Although factual Time-Sensitive Que…
Differentially Private Federated Clustering with Random Rebalancing
Xiyuan Yang, Shengyuan Hu, Soyeon Kim +1
Federated clustering aims to group similar clients into clusters and produce one model for each cluster. Such a personalization approach typically improves model performance compar…
NEXT-EVAL: Next Evaluation of Traditional and LLM Web Data Record Extraction
Soyeon Kim, Namhee Kim, Yeonwoo Jeong
Effective evaluation of web data record extraction methods is crucial, yet hampered by static, domain-specific benchmarks and opaque scoring practices. This makes fair comparison b…
PFGuard: A Generative Framework with Privacy and Fairness Safeguards
Soyeon Kim, Yuji Roh, Geon Heo +1
Generative models must ensure both privacy and fairness for Trustworthy AI. While these goals have been pursued separately, recent studies propose to combine existing privacy and f…
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
Jio Oh, Soyeon Kim, Junseok Seo +4
Large language models (LLMs) have achieved unprecedented performances in various applications, yet evaluating them is still challenging. Existing benchmarks are either manually con…