3 papers
cs.LG2026
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
Tong Liu, Cheng Qian, Matej Cief +4
Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies tool-calling along…
cs.LG2024
Cross-Validated Off-Policy Evaluation
Matej Cief, Branislav Kveton, Michal Kompan
We study estimator selection and hyper-parameter tuning in off-policy evaluation. Although cross-validation is the most popular method for model selection in supervised learning, o…
cs.LG2024
Pessimistic Off-Policy Optimization for Learning to Rank
Matej Cief, Branislav Kveton, Michal Kompan
Off-policy learning is a framework for optimizing policies without deploying them, using data collected by another policy. In recommender systems, this is especially challenging du…