Empirical Metrics & Validation¶
This document specifies the quantitative metrics and automated validation checks employed by Episteme to evaluate the technical quality, structural fidelity, and retrieval utility of the constructed knowledge and theory graphs.
Tip
Theory Graph & Epistemic Metrics (5 Pillars):
This document focuses specifically on empirical construction metrics (Phases 1 through 5). Formal epistemic
metrics (Gerhard Schurz's empirical creativity, unification index, and homogeneity against the tacking paradox;
Imre Lakatos's degeneration index; Dung's abstract argumentation semantics; and Paul Thagard's explanatory coherence)
are formalized in the Theory Metrics Overview & Taxonomy and developed in the
companion package epistemetrics.
Metric Taxonomy & Architectural Scope¶
Evaluation metrics in Episteme are organized into five empirical categories:
graph TD
M1["1. Graph Structural Metrics (GM-GBS, OEP)"]
M2["2. Entity & Relation Quality (NER, Linking, Extraction)"]
M3["3. Argument Mining Metrics (ADU, ACC, ARC)"]
M4["4. Extrinsic Utility (MRR, Hits@k, nDCG, MAP)"]
M5["5. Automated Graph Integrity (Neo4j Constraints & Provenance)"]
M1 --> EVAL["Evaluation Report (report_template.md)"]
M2 --> EVAL
M3 --> EVAL
M4 --> EVAL
M5 --> EVAL
Graph Structural Metrics (Intrinsic Evaluation)¶
Graph structural metrics quantify the topological and semantic alignment between the LLM-generated graph and a Gold Standard, replacing rigid string equality with contextual embedding similarity.
| Metric | Definition & Formula | Target Pipeline Phase | Implementation / Tooling |
|---|---|---|---|
| GM-GBS (Graph BERTScore) | Soft-matching F1 over relation labels using cosine similarity of contextual embeddings above threshold \(\tau = 0.95\): \(\text{Soft-F1} = \frac{2 \cdot P_{\text{soft}} \cdot R_{\text{soft}}}{P_{\text{soft}} + R_{\text{soft}}}\) |
Phase 2, Phase 3, Phase 5 | gm_gbs.py |
| Hallucination Rate (\(HR\)) | Fraction of predicted edges with no semantic match in the gold standard: $HR = \frac{ |
E_{\text{pred}} | - |
| Omission Rate (\(OR\)) | Fraction of gold-standard edges omitted by the pipeline: $OR = \frac{ |
E_{\text{gold}} | - |
| Graph Edit Distance (GED) | Minimum cost sequence of node/edge insertions, deletions, and substitutions required to transform \(G_{\text{pred}}\) into \(G_{\text{gold}}\). | Phase 3b, Phase 5 | NetworkX / Scipy |
Entity & Relationship Quality Metrics¶
These metrics evaluate the accuracy of entity discovery, canonical resolution, and relationship extraction against annotated ground truth.
| Task / Quality Dimension | Specific Metric | Definition & Diagnostic Value | Target Phase |
|---|---|---|---|
| Entity Extraction (NER) | Per-Type Precision, Recall, F1 | Evaluates boundary detection and coarse/fine typing (e.g., Person, Concept, Work, School). | Phase 2 |
| Entity Linking / Resolution | Pairwise Accuracy | Correctness of SAME_AS equivalence decisions between entity mentions. |
Phase 3, Phase 5 |
| Cluster Purity & Completeness | Purity measures absence of alien entities in a cluster; completeness measures absence of missed co-references. | Phase 5 | |
| False Merge Rate | Fraction of distinct theoretical concepts erroneously merged into a single node. | Phase 5 | |
| False Split Rate | Fraction of coreferent mentions left unmerged as duplicate nodes. | Phase 5 | |
| Local Relation Extraction | Relation Precision / Recall | Accuracy of extracted triples within individual chunks against ground-truth relations. | Phase 2 |
| Evidence Grounding Ratio | Proportion of extracted relations backed by valid verbatim source character spans. | Phase 2 | |
| Global Relation Discovery | Candidate Recall@k | Proportion of gold-standard cross-chunk relations present in top-\(k\) retrieved candidates. | Phase 3 |
| Reranking MRR | Rank position of valid relation candidates after cross-encoder reranking. | Phase 3 |
Argument Mining Metrics¶
Argument mining metrics evaluate the extraction and structuring of argumentative discourse units (ADUs) and defeasible reasoning patterns.
| Task Level | Metric | Operational Measurement | Target Phase |
|---|---|---|---|
| ADU Segmentation | Span Exact / Partial Match (IoU) | Intersection-over-Union (\(\text{IoU} \ge 0.70\)) of character spans between predicted and gold discourse units. | Phase 4 |
| Component Classification (ACC) | Macro-Averaged F1 | Multi-class classification accuracy across argument roles: Premise, Claim, Major Claim, Axiom. | Phase 4 |
| Relation Classification (ARC) | Support / Attack F1 | Binary and multi-class classification for argumentative links (SUPPORTS, REFUTES, QUALIFIES). |
Phase 4, Phase 5 |
| Structural Consistency | Verification that intra-argument derivation chains form valid Directed Acyclic Graphs (DAGs). | Phase 4, Phase 6 |
Extrinsic Information Retrieval Metrics¶
Extrinsic metrics quantify the practical utility of the Knowledge Graph when queried by downstream tasks (Semantic Search, Question Answering, and Literature Survey).
| Metric | Mathematical Formulation | Interpretation |
|---|---|---|
| Mean Reciprocal Rank (MRR) | $\text{MRR} = \frac{1}{ | Q |
| Hits@k | $\text{Hits}@k = \frac{1}{ | Q |
| nDCG@k | \(\text{nDCG}@k = \frac{\text{DCG}@k}{\text{IDCG}@k}, \quad \text{DCG}@k = \sum_{i=1}^k \frac{\text{rel}_i}{\log_2(i + 1)}\) | Evaluates graded relevance ranking, penalizing models that place highly relevant theoretical constructs below marginal results. |
| Mean Average Precision (MAP) | $\text{MAP} = \frac{1}{ | Q |
Reference-Free Evaluation & Graph Integrity¶
Reference-Free Metrics (LLM-as-a-Judge)¶
In accordance with our Evaluation Methodology, reference-free evaluation using LLMs as judges is strictly restricted to Layer 1 & 2 extraction quality, and must be calibrated against human review baselines:
| Metric | Prompt Assessment Goal | Calibration Requirement |
|---|---|---|
| Faithfulness | Verifies whether an extracted entity or triple is fully supported by the source text chunk (detects hallucinations). | Correlation \(\ge 0.80\) with human expert judgments on review samples. |
| Comprehensiveness | Verifies whether all major theoretical concepts in a passage were extracted (detects omissions). | Adjudicated against human expert review sets. |
Graph Integrity & Constraint Verification (Neo4j)¶
Automated graph validation ensures that the constructed graph satisfies formal relational and data integrity constraints:
- Schema Compliance: Every node and relationship matches the taxonomy defined in
pipeline/schema/. - Provenance Coverage: Percentage of nodes and edges possessing valid
source_doc_id,chunk_id, and character offsets. - Uniqueness Constraints: Absence of duplicate canonical entity IDs or duplicate directed edges with identical predicates.
- Orphan Detection: Verification that theoretical claims and premises maintain logical connections to surrounding discourse.
Related Documentation¶
- Evaluation Methodology & Scaffolding: Evaluation Methodology
- Evaluation Harness Architecture: Evaluation Harness
- Dataset Portfolio & Strategies: Datasets Strategy
- Epistemic & Theory Graph Metrics (5 Pillars): Theory Metrics Overview
- Formal Graph Model: Formal Graph Schema (TheoryNet)
- Observability Integration: ADR 0012: Langfuse Integration