4 papers
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
Jitian Zhao, Changho Shin, Tzu-Heng Huang +2
LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges prov…
Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
Jitian Zhao, Chenghui Li, Frederic Sala +1
Concept-based approaches, which aim to identify human-understandable concepts within a model's internal representations, are a promising method for interpreting embeddings from dee…
MoRe Fine-Tuning with 10x Fewer Parameters
Wenxuan Tan, Nicholas Roberts, Tzu-Heng Huang +5
Parameter-efficient fine-tuning (PEFT) techniques have unlocked the potential to cheaply and easily specialize large pretrained models. However, the most prominent approaches, like…
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
Changho Shin, Jitian Zhao, Sonia Cromp +2
Popular zero-shot models suffer due to artifacts inherited from pretraining. One particularly detrimental issue, caused by unbalanced web-scale pretraining data, is mismatched labe…