API security 1distillation attacks 1large language models 1output perturbation defenses 1threat modeling 1
From the 1 of 4 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
Daniel Yang, Samuel Stante, Florian Redhardt +5
Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncert…
cs.LG2025
Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
Lena Libon, Meghana Bhange, Rushabh Solanki +2
The current era of AI development places a heavy emphasis on training large models on increasingly scaled-up datasets. This paradigm has catalyzed entirely new product categories,…