3 papers
cs.LG2025
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
Saeed Mohammadzadeh, Erfan Hamdi, Joel Shor +1
As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical mod…
cs.HC2025
From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems
Marcos Ortiz, Justin Hill, Collin Overbay +4
Agentic AI systems capable of generating full-stack web applications from natural language prompts ("prompt- to-app") represent a significant shift in software development. However…
q-bio.QM2025
Cross-Domain Transfer of Depression Voice Biomarkers Depends on the Outcome Instrument: Leakage-Controlled Cross-Sectional Evaluation Study
Rachel L. Wiley, James Schwoebel, Matias Caccia +4
Whether voice biomarkers of depression generalize across clinical settings is largely untested. Generalization is usually framed as a question about populations. It is also a quest…