3 citations · 3 across the 1 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
Sai Kolasani, Maxim Saplin, Nicholas Crispino +5
We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extend…
cs.AI2025
Optimizing Model Selection for Compound AI Systems
Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4
Compound AI systems that combine multiple LLM calls, such as self-refine and multi-agent-debate, achieve strong performance on many AI tasks. We address a core question in optimizi…