Measuring Knowledge in AI Systems: Automated domain specific Knowledge Graph Construction from Scientific Literature

What is knowledge, and how do we measure it? Extracted knowledge is more than a true statement; it carries justification, provenance, and a context of use. When an expert builds a knowledge graph from scientific literature, every node and edge reflects a justified interpretation: they know why a relation holds, which paper supports it, and under what conditions it applies. This raises a fundamental challenge for AI-based information retrieval systems. Retrieval-augmented generation (RAG) can now produce knowledge graphs that resemble those assembled by hand, yet these graphs are generated through statistical pattern recognition rather than contextual understanding of the scientific processes they describe. The central question of this work is therefore: how can we measure whether such systems are epistemic agents or only convincing imitators of one?

To explore this problem, we developed OntoGeoRAG, an open-source framework that combines document retrieval, LLM-based information extraction, and entity–relation normalization to transform scientific PDFs into provenance-aware semantic knowledge graphs. Using mass-transport deposits in subsurface geophysics as a challenging test case, the framework achieves 77% recovery against an expert-curated benchmark. Rather than treating the remaining 23% as simple error, we interpret these failures as evidence of the current epistemic limitations of retrieval-based AI systems. Although demonstrated in geophysics, the methodology is domain-agnostic and transferable to any scientific field with structured terminology and complex conceptual relationships.


Knowledge locked in scientific PDFs remains largely inaccessible to computation. Domain experts spend significant time manually reading, extracting, and encoding entity–relation triples into knowledge graphs, a process that is slow, inconsistent, and difficult to scale. OntoGeoRAG addresses this bottleneck with a fully open-source pipeline that automates the transformation from PDF documents to validated, quarriable knowledge graphs while preserving traceability to the original scientific evidence.

The pipeline consists of four stages. First, schema-driven query generation derives targeted retrieval questions from an expert-defined ontology. Second, the retrieval module segments and indexes scientific documents before ranking the most relevant passages using keyword retrieval followed by neural re-ranking. Third, an open-source large language model extracts entities and relations from the retrieved passages to produce candidate semantic triples. Finally, each extracted triple is verified against its source passage, linking every relation back to the original text in order to construct a reliable and provenance-aware knowledge graph.

The framework was evaluated on a corpus of 41 geological papers focused on mass-transport deposits, a particularly difficult testbed because the terminology is ambiguous, the entities are heterogeneous, and the relations often depend on implicit geological reasoning. Compared to an expert-curated benchmark graph, OntoGeoRAG successfully recovered 77% of expert-defined relations.

Importantly, the remaining 23% of missing or incorrect relations reveal critical limitations of current RAG-based approaches. These failures are not random extraction errors; they frequently occur in contexts where scientific meaning depends on implicit reasoning, sparse evidence, long-range contextual dependencies, causal interpretation, or domain-specific conceptual understanding not explicitly stated within nearby text passages. In many cases, the system retrieves statistically associated information without reconstructing the conceptual logic that allows an expert to justify why a relation is scientifically valid.

Our analysis shows that retrieval quality is the key bottleneck in domain-specific knowledge extraction using open-source LLMs. When relevant contextual evidence is absent, fragmented, or weakly ranked during retrieval, the downstream language model cannot recover scientifically meaningful relations. This observation suggests that current RAG systems remain fundamentally pattern-recognition machines rather than systems capable of genuine contextual understanding.

Although evaluated in subsurface geophysics, the framework itself imposes no geological assumptions and is transferable to any scientific domain with a defined conceptual schema, including biology, medicine, environmental science, and engineering. More broadly, this work reframes automated knowledge graph construction as an epistemological problem: not simply whether AI systems can retrieve information, but whether they can construct, justify, and validate scientific knowledge itself.

Feryal Batoul TALBI

PhD researcher with a background in computer science, currently a Marie Skłodowska-Curie fellow at Sorbonne Université (LIP6) and IFP Energies nouvelles. My work focuses on applied AI and machine learning, with particular interest in retrieval-augmented generation, knowledge graphs, large language models, and unsupervised deep learning.