Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
RLNVR: Reinforcement Learning from Non-Verified Real-World Rewards
Rohit Krishnan, Jon Evans
This paper introduces RLNVR (Reinforcement Learning from Non-Verified Rewards), a framework for training language models using noisy, real-world feedback signals without requiring…
cs.AI2025
Deep Research Bench: Evaluating AI Web Research Agents
FutureSearch, :, Nikos I. Bosse +7
Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the…