3 papers
cs.LG2025
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
Bernd Bohnet, Pierre-Alexandre Kamienny, Hanie Sedghi +7
We demonstrate an approach for LLMs to critique their \emph{own} answers with the goal of enhancing their performance that leads to significant improvements over established planni…
cs.CY2025
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
Irina Jurenka, Markus Kunesch, Kevin R. McKee +71
A major challenge facing the world is the provision of equitable and universal access to quality education. Recent advances in generative AI (gen AI) have created excitement about…
cs.CY2025
Evaluating Gemini in an arena for learning
LearnLM Team, Abhinit Modi, Aditya Srikanth Veerubhotla +34
Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-t…