From the 1 of 12 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry
Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3
Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design problems to be difficult for humans or c…
cs.AI2025
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
Tian Qin, Core Francisco Park, Mujin Kwun +5
Mathematical reasoning tasks have become prominent benchmarks for assessing the reasoning capabilities of LLMs, especially with reinforcement learning (RL) methods such as GRPO sho…