3 papers
cs.LG2026
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
Saeed Mohammadzadeh, Erfan Hamdi, Joel Shor +1
As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical mod…
cs.HC2026
From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems
Marcos Ortiz, Justin Hill, Collin Overbay +4
Agentic AI systems capable of generating full-stack web applications from natural language prompts ("prompt- to-app") represent a significant shift in software development. However…
q-bio.QM2025
Voice biomarkers of perinatal depression: cross-sectional nationwide pilot study report
Rachel L. Wiley, Jim Schwoebel, Joel Shor +5
Perinatal depression (PND) affects 1 in 5 mothers, with 85% lacking support. Digital health tools offer early identification and prevention, potentially reducing PND risk by over 5…