1 paper · 1 filter
Hyunjun Kim, Sejong Kim
We introduce MacroBench, a code-first benchmark that evaluates whether LLMs can synthesize reusable browser-automation programs (macros) from natural-language goals by reading HTML…