Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Ali Ansari, Haoran Sun, Andy Zeyi Liu +48
Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still strugg…
cs.AI2026
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
Ali Ansari, Yasmin Mohammadi, Farnoush Nili +3
Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiti…