3 papers
cs.LG2026
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
Pin Qian, Su Wang, Yihang Chen +5
Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in…
cs.AI2026
Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks
Jihan Yao, Gantavya Bhatt, Arnav Das +16
We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the f…
cs.DC2025
Efficient Federated Learning with Heterogeneous Data and Adaptive Dropout
Ji Liu, Beichen Ma, Qiaolin Yu +7
Federated Learning (FL) is a promising distributed machine learning approach that enables collaborative training of a global model using multiple edge devices. The data distributed…