Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
SPEAR: Code-Augmented Agentic Prompt Optimization
Mengyin Lu, Cong Feng, Huimin Han +6
Automatic prompt engineering (APE) rewrites prompts to improve downstream task performance, but existing APE loops treat the optimizer itself as a fixed pipeline. We port the code-…
cs.CL2026
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
Gaurav Srivastava, Aafiya Hussain, Zhenyu Bi +5
Evaluating language models fairly is increasingly difficult as static benchmarks risk contamination by training data, obscuring whether models truly reason or recall. We introduce…
cs.CL2025
DEBATE, TRAIN, EVOLVE: Self Evolution of Language Model Reasoning
Gaurav Srivastava, Zhenyu Bi, Meng Lu +1
Large language models (LLMs) have improved significantly in their reasoning through extensive training on massive datasets. However, relying solely on additional data for improveme…