2 papers
cs.LG2026
STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6
Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majorit…
cs.CR2025
DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge
Asmita Mohanty, Gezheng Kang, Lei Gao +1
Large Language Models (LLMs) have demonstrated strong performance across diverse tasks, but fine-tuning them typically relies on cloud-based, centralized infrastructures. This requ…