Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
Noy Sternlicht, Ariel Gera, Roy Bar-Haim +2
We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges. Evaluating debate speeches requires a deep understanding of the speech at multi…
cs.CL2024
Conversational Prompt Engineering
Liat Ein-Dor, Orith Toledo-Ronen, Artem Spector +5
Prompts are how humans communicate with LLMs. Informative prompts are essential for guiding LLMs to produce the desired output. However, prompt engineering is often tedious and tim…