6 papers
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
Aakriti Agrawal, Mucong Ding, Zora Che +6
With Large Language Models (LLMs) rapidly approaching and potentially surpassing human-level performance, it has become imperative to develop approaches capable of effectively supe…
Effort-aware Fairness: Incorporating a Philosophy-informed, Human-centered Notion of Effort into Algorithmic Fairness Metrics
Tin Trung Nguyen, Jiannan Xu, Zora Che +6
Although popularized AI fairness metrics, e.g., demographic parity, have uncovered bias in AI-assisted decision-making outcomes, they do not consider how much effort one has spent…
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Xiyao Wang, Zhengyuan Yang, Linjie Li +6
Despite significant advancements in vision-language models (VLMs), there lacks effective approaches to enhance response quality by scaling inference-time computation. This capabili…
SAFLEX: Self-Adaptive Augmentation via Feature Label Extrapolation
Mucong Ding, Bang An, Yuancheng Xu +2
Data augmentation, a cornerstone technique in deep learning, is crucial in enhancing model performance, especially with scarce labeled data. While traditional techniques are effect…
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
Aakriti Agrawal, Mucong Ding, Zora Che +6
With Large Language Models (LLMs) rapidly approaching and potentially surpassing human-level performance, it has become imperative to develop approaches capable of effectively supe…
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
Mucong Ding, Chenghao Deng, Jocelyn Choo +8
While generalization over tasks from easy to hard is crucial to profile language models (LLMs), the datasets with fine-grained difficulty annotations for each problem across a broa…