Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Efficient Agent Evaluation via Diversity-Guided User Simulation
Itay Nakash, George Kour, Ateret Anaby-Tavor
Large language models (LLMs) are increasingly deployed as customer-facing agents, yet evaluating their reliability remains challenging due to stochastic, multi-turn interactions. C…
cs.AI2025
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
George Kour, Itay Nakash, Ateret Anaby-Tavor +1
As Large Language Models (LLMs) become deeply integrated into human life and increasingly influence decision-making, it's crucial to evaluate whether and to what extent they exhibi…