natural language processing

Pezego-HITL: A policy-grounded large language model architecture for agricultural extension in Ghana

arXiv:2607.13934

summary

The paper presents Pezego-HITL, a policy‑grounded large language model system for agricultural extension in Ghana, and introduces the P‑EVAL framework to evaluate safety, usefulness, latency, and expert workload in decision‑support queries.

Abstract

Large language models are increasingly deployed in agricultural decision-support settings, yet high-stakes crop protection in smallholder agriculture requires more than output-quality benchmarks. Over a two-year design and evaluation programme, we formalise policy-constrained large language model assessment as an adaptive compute allocation problem that jointly captures safety compliance, helpfulness, operational latency, and expert supervision workload. We introduce P-EVAL (Policy-grounded Expert-calibrated VALidation protocol), a unified evaluation framework for policy-grounded decision support, evaluating the architecture on a simulated field query database consisting of 1,240 cases. The protocol is instantiated on the Pezego advisory architecture (Pezego-HITL) and evaluated in Ghana. Following offline judge calibration against gold-standard human expert decisions (), we evaluate the architectural performance under simulated query workloads. Under P-EVAL, our memory-routed architecture improves the Policy Alignment Rate (PAR) to 0.94 and the Agronomic Utility Rate (AUR) to 0.95, while reducing P95 latency by 55% (from 28.6s to 12.9s) through a 59.6% cache reuse ratio. We also demonstrate generalisability using the open-source \texttt{Qwen3.5-9B-DeepSeek-V4-Flash} model, achieving a PAR of 0.86 and a 54.5% latency reduction (to 10.2s). To evaluate practical utility and socio-technical integration, we administer detailed questionnaires to Ghanaian Extension Services Officers () and smallholder farmers (). Taken together, this work demonstrates how policy-grounded structured retrieval-augmented generation with validated-memory routing makes safety-utility-latency trade-offs explicit, offering a scalable template for trustworthy AI-driven extension in smallholder farming systems.

Topics & keywords

#agricultural extension#large language models#policy alignment#human-in-the-loop#retrieval‑augmented generationP‑EVALpolicy‑grounded validationmemory‑routed architectureQwen3.5-9B-DeepSeek-V4-Flashpolicy alignment rateagronomic utility rate