Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Bingyang Ye, Shan Chen, Jingxuan Tu +4
Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientif…
cs.CL2025
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
Jirui Qi, Shan Chen, Zidi Xiong +3
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studi…