Knowledge-graph-augmented retrieval for RAG systems — entity extraction, co-occurrence graphs, and graph-based document retrieval via Neo4j.
pip install ragleap-graph
Quickstart
from ragleap_graph import GraphConfig, GraphIndex
graph = GraphIndex(config=GraphConfig(
uri="bolt://localhost:7687",
user="neo4j",
password="...",
))
graph.upsert_document(
document_id="doc-1",
title="Q3 Report",
chunks=[{"text": "Acme Corp reported strong Q3 revenue growth."}],
)
docs = graph.find_documents_by_entities(["Acme Corp"])
related = graph.search_related_entities(["Acme Corp"], max_depth=2)
Current: v0.9.0 · 108 tests (92 passed without live credentials, 16 skipped)
Architecture
LLM-based extraction and dedup (v0.2.0+)
The default entity extraction is regex/heuristic-based (fast, free, zero
dependencies). For messier input — e.g. inconsistent capitalization like
"Acme Corp" vs "ACME Corp." — LLM-based extraction and dedup produce
cleaner graphs. Requires the llm extra: pip install ragleap-graph[llm]
from ragleap.generation import ProviderConfig
from ragleap_graph import GraphConfig, GraphIndex, ExtractionConfig
graph = GraphIndex(
config=GraphConfig(uri="bolt://localhost:7687", user="neo4j", password="..."),
extraction=ExtractionConfig(
method="llm",
provider=ProviderConfig(provider="gemini", api_key="...", model="gemini-3.6-flash"),
dedup_enabled=True,
),
)
Note: EntityDeduplicator merges spelling variants of an already-extracted
name; it does not fix fragmentation caused by the regex extractor splitting
one real-world entity into multiple candidates in the first place — see
CHANGELOG.md for a real, measured example of this and how method="llm"
avoids it at the source.
Cross-chunk relation extraction (v0.8.0+)
Relation extraction normally runs per-chunk, so a relation whose evidence
spans two chunks — e.g. an entity named in chunk 1, referred to only by
pronoun ("it", "the company") in a later chunk — can be missed. Enable
cross_chunk_relations=True for one additional pass over the full
document using every entity found across all chunks:
extraction=ExtractionConfig(
method="llm",
provider=ProviderConfig(provider="gemini", api_key="...", model="gemini-3.6-flash"),
extract_relations=True,
cross_chunk_relations=True,
)
Known limitation: this depends on the provider's reasoning ability to
resolve the reference correctly. Live-verified working with Gemini;
small local models (e.g. qwen2.5:0.5b) were found, via live testing,
to produce an incorrect relation rather than none on this task — not
recommended for this feature without independently verifying its output.
Ontology cross-validation (v0.9.0+)
Constrain which relation_type values are valid between which
entity_type pairs. A relation violating the ontology is dropped (with
a WARNING logged) rather than written to Neo4j; relation_type values
not listed remain unconstrained. Requires entity_types= to also be set:
extraction=ExtractionConfig(
method="llm",
provider=ProviderConfig(provider="gemini", api_key="...", model="gemini-3.6-flash"),
extract_relations=True,
entity_types=["ORGANIZATION", "PERSON"],
relation_ontology={"FOUNDED_BY": (["ORGANIZATION"], ["PERSON"])},
)
Operations
ragleap-graph is a knowledge-graph retrieval library, not a database
administration tool. Backup and restore of the underlying Neo4j database
are explicitly the operator's responsibility, not something this library
wraps or automates -- see
docs/operations/backup-and-restore.md
for a concrete, tested procedure using Neo4j's native neo4j-admin
tooling, and
docs/adr/0001-backup-restore-ownership.md
for the reasoning behind this scope decision.
Status
v0.9.0. Ported from a real production GraphService, adapted for standalone open-source use — see HANDOFF.md for the full design history. ragleap-rag >=0.12.0 is an optional dependency, required only for method="llm".
License
MIT