2 papers
cs.AI2026
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
Linghao Feng, Yinqian Sun, Dongqi Liang +8
Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning a…
cs.AI2026
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
Jianghao Lin, Yuanyuan Shi, Xin Peng +10
Large language models (LLMs) excel at function calling, but inference scaling has been explored mainly for unstructured generation. We propose an inference-scaling framework for st…