2 papers
cs.LG2025
Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
Petr Spelda, Vit Stritecky
Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before tas…
cs.CR2025
Security practices in AI development
Petr Spelda, Vit Stritecky
What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment a…