benchmark 1counterfactual evaluation 1inference regularization 1language grounding 1robot control 1vision-language models 1
From the 1 of 6 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
Zheng Chu, Xiao Wang, Jack Hong +11
Large language models are transitioning from generalpurpose knowledge engines to realworld problem solvers, yet optimizing them for deep search tasks remains challenging. The centr…
cs.AI2026
Detecting RLVR Training Data via Structural Convergence of Reasoning
Hongbo Zhang, Yue Yang, Jianhao Yan +2
Reinforcement learning with verifiable rewards (RLVR) is central to training modern reasoning models, but the undisclosed training data raises concerns about benchmark contaminatio…