11 papers
What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence Interventions
Wenzhang Du
A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair. We introduce Pair-ID, an offline audit that holds one query…
Optimistic Feasible Search for Closed-Loop Fair Threshold Decision-Making
Wenzhang Du
Closed-loop decision-making systems (e.g., lending, screening, or recidivism risk assessment) often operate under fairness and service constraints while inducing feedback effects:…
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
Wenzhang Du
Subjective mean opinion scores (MOS) remain the de-facto target for non-intrusive speech and singing quality assessment. However, MOS is a scalar that collapses heterogeneous user…
Contract-Governed Training for Earth Observation: Observed Service Agreement Graphs and Coverage-Accuracy Trade-offs
Wenzhang Du
Earth observation (EO) models are frequently trained under implicit sampling policies that optimize global accuracy but provide no explicit guarantees on who (which regions, classe…
Retrieval-Augmented Memory for Online Learning
Wenzhang Du
Retrieval-augmented models couple parametric predictors with non-parametric memories, but their use in streaming supervised learning with concept drift is not well understood. We s…
City-Conditioned Memory for Multi-City Traffic and Mobility Forecasting
Wenzhang Du
Deploying spatio-temporal forecasting models across many cities is difficult: traffic networks differ in size and topology, data availability can vary by orders of magnitude, and n…