Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Propose, Solve, Verify: Self-Play Through Formal Verification
Alex Wilf, Pranjal Aggarwal, Bryan Parno +4
Training models through self-play alone (without any human data) has been a longstanding goal in AI, but its effectiveness for training large language models remains unclear, parti…
cs.AI2025
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
Gyeongwon James Kim, Alex Wilf, Louis-Philippe Morency +1
Recent progress in autonomous code generation has fueled excitement around AI agents capable of accelerating scientific discovery by running experiments. However, there is currentl…