2 papers
cs.CR2026
Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models
Weidi Luo, Xiaofei Wen, Tenghao Huang +5
Large language models (LLMs) are increasingly deployed for everyday tasks, including food preparation and health-related guidance. However, food safety remains a high-stakes domain…
cs.CL2025
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
Yixin Cao, Shibo Hong, Xinze Li +24
Large Language Models (LLMs) are advancing at an amazing speed and have become indispensable across academia, industry, and daily applications. To keep pace with the status quo, th…