Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
Shuangshuang Ying, Zheyu Wang, Yunjian Peng +16
Despite strong performance on existing benchmarks, it remains unclear whether large language models can reason over genuinely novel scientific information. Most evaluations score e…
cs.AI2025
OAgents: An Empirical Study of Building Effective Agents
He Zhu, Tianrui Qin, King Zhu +21
Recently, Agentic AI has become an increasingly popular research field. However, we argue that current agent research practices lack standardization and scientific rigor, making it…