activity
20242026
collaborators

5 papers

cs.CL2026

Efficient semantic uncertainty quantification in language models via diversity-steered sampling

Ji Won Park, Kyunghyun Cho

Accurately estimating semantic aleatoric and epistemic uncertainties in large language models (LLMs) is particularly challenging in free-form question answering (QA), where obtaini…

cs.LG2025

Semiparametric conformal prediction

Ji Won Park, Robert Tibshirani, Kyunghyun Cho

Many risk-sensitive applications require well-calibrated prediction sets over multiple, potentially correlated target variables, for which the prediction algorithm may report corre…

cs.LG2025

Supervised Contrastive Block Disentanglement

Taro Makino, Ji Won Park, Natasa Tagasovska +11

Real-world datasets often combine data collected under different experimental conditions. This yields larger datasets, but also introduces spurious correlations that make it diffic…

cs.LG2024

Concept Bottleneck Language Models For protein design

Aya Abdelsalam Ismail, Tuomas Oikarinen, Amy Wang +8

We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our arc…

cs.LG2024

Generalizing to any diverse distribution: uniformity, gentle finetuning and rebalancing

Andreas Loukas, Karolis Martinkus, Ed Wagstaff +1

As training datasets grow larger, we aspire to develop models that generalize well to any diverse test distribution, even if the latter deviates significantly from the training dat…