11 papers
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
Riccardo Terrenzi, Serkan Ayvaz
Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. We study six metadata-generation settings…
PIPER: Content-Based Table Search via profiling and LLM-Generated Pseudoqueries
Riccardo Terrenzi, Matteo Falconi, Serkan Ayvaz +1
The rapid growth of tabular datasets in data lakes, data spaces, and open data portals makes effective dataset search essential for reuse and analysis. Existing search systems rely…
Why Neighborhoods Matter: Traversal Context and Provenance in Agentic GraphRAG
Riccardo Terrenzi, Maximilian von Zastrow, Serkan Ayvaz
Retrieval-Augmented Generation can improve factuality by grounding answers in external evidence, but Agentic GraphRAG complicates what it means for citations to be faithful. In the…
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed un…
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
Phongsakon Mark Konrad, Tim Lukas Adam, Ane Cathrine Holst Merrild +4
AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model internals can be explained. Thi…
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Safe fine-tuning defenses are often endorsed on the basis of a held-out gap reduction, but the same reduction can come from sampling noise, subject artifacts, capability loss, or a…