3 papers
cs.AI2025
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
Jirui Yang, Hengqi Guo, Zhihui Lu +6
Large language models often face a three-way trade-off among detection accuracy, inference latency, and deployment cost when used in real-world safety-sensitive applications. This…
cs.CL2025
Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark
Hua Zhou, Bing Ma, Yufei Zhang +1
This paper comprehensively elaborates on the construction methodology, multi-dimensional evaluation system, and underlying design philosophy of CUFEInse v1.0. Adhering to the princ…
cs.HC2025
Generative AI alone may not be enough: Evaluating AI Support for Learning Mathematical Proof
Eason Chen, Sophia Judicke, Kayla Beigh +16
We evaluate the effectiveness of LLM-Tutor, a large language model (LLM)-powered tutoring system that combines an AI-based proof-review tutor for real-time feedback on proof-writin…