2 papers
cs.CL2026
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
Sara Vera MarjanoviÄ, Arkil Patel, Vaibhav Adlakha +14
Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creat…
cs.LG2025
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade +7
Web agents enable users to perform tasks on web browsers through natural language interaction. Evaluating web agents trajectories is an important problem, since it helps us determi…