activity
20242026
collaborators

14 papers

cs.CL2026

GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski

We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canon…

cs.CL2026

GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski

Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model…

cs.CL2026

GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge

Yujia Hu, Tuan-Phong Nguyen, Shrestha Ghosh +2

Language models are powerful artifacts, yet their factual knowledge is still poorly understood, and inaccessible to ad-hoc browsing and scalable statistical analysis. This demonstr…

cs.CL2026

LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic Knowledge at Scale

Muhammed Saeed, Simon Razniewski

Benchmarks like MMLU suggest flagship language models approach factuality saturation above 90\%. \emph{LLMpedia} shows this picture is incomplete. We materialize 1.3M encyc…

cs.CL2026

Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval

Valentin Knappich, Anna Hätty, Simon Razniewski +1

Novelty assessment is a critical yet complex task in the examination process for patent acceptance, requiring examiners to determine whether an invention is disclosed in a prior ar…

cs.CL2026

Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness

Luca Giordano, Simon Razniewski

Large Language Models (LLMs) encode substantial factual knowledge, yet measuring and systematizing this knowledge remains challenging. Converting it into structured format, for exa…