Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems
Rebecca L. Johnson
In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: benchmarks do not just measure,…
cs.AI2026
How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge? A Software Engineering Framework
Choro Ulan uulu, Mikhail Kulyabin, Iris Fuhrmann +6
Critical domain knowledge typically resides with few experts, creating organizational bottlenecks in scalability and decision-making. Non-experts struggle to create effective visua…