3 papers
cs.CL2026
Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge
Zijie Wang, Eduardo Blanco
LLM-as-a-Judge is a scalable alternative to human evaluation, yet existing rubric-based methods rely on human-annotated data such as reference answers or expert-crafted rubrics. We…
cs.AI2026
Error Taxonomy-Guided Prompt Optimization
Mayank Singh, Vikas Yadav, Eduardo Blanco
Automatic Prompt Optimization (APO) is a powerful approach for extracting performance from large language models without modifying their weights. Many existing methods rely on tria…
cs.AI2025
Grammar Search for Multi-Agent Systems
Mayank Singh, Vikas Yadav, Shiva Krishna Reddy Malay +4
Automatic search for Multi-Agent Systems has recently emerged as a key focus in agentic AI research. Several prior approaches have relied on LLM-based free-form search over the cod…