2 papers
cs.AI2026
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
Shuangshuang Ying, Zheyu Wang, Yunjian Peng +16
Despite strong performance on existing benchmarks, it remains unclear whether large language models can reason over genuinely novel scientific information. Most evaluations score e…
cs.CL2026
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
Shuangshuang Ying, Yunwen Li, Xingwei Qu +21
Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We intr…