The mental model
You loaded an ontology and some data. Behind the asserted triples there is a chain of consequences: subclass and subproperty closures, equivalent classes, inverse and transitive properties, owl:sameAs merges.
pgrdf.materialize(graph_id) computes those consequences and stores them as real triples in the same graph, flagged as inferred. The triples you loaded stay flagged as asserted.
After materialization, SPARQL queries and SHACL validation see asserted and inferred triples as one flat set. Your application doesn't need to know which is which. When you do need the split, graph_inventory() counts each kind, and export_graph returns the asserted triples only.
What changes after a materialize call
| Asserted | Also present after pgrdf.materialize(g) |
|---|---|
ex:alice a ex:Engineer, ex:Engineer rdfs:subClassOf ex:Person, ex:Person rdfs:subClassOf ex:Agent | ex:alice a ex:Person, ex:alice a ex:Agent |
ex:alice ex:knows ex:bob, ex:knows owl:inverseOf ex:knownBy | ex:bob ex:knownBy ex:alice |
ex:ali owl:sameAs ex:alice, ex:ali ex:worksAt ex:acme | ex:alice ex:worksAt ex:acme |
The 'owl-rl' profile also adds a few axiomatic triples, such as typing each resource as owl:Thing.
The return value reports how many triples were read, how many were inferred, and how many inferred triples from the previous run were replaced.
A snapshot, refreshed on demand
Inferred triples reflect the graph as it was when materialize ran. Loading or deleting asserted triples afterwards doesn't update them. Run materialize again to bring them up to date. It replaces the old inferred set rather than adding to it.
graph_inventory() shows whether a graph needs a refresh. See Is the materialization current?
Cost, and why size matters
Materialization is forward chaining. The reasoner computes the whole closure in memory, then pgRDF writes the triples that weren't already asserted. Runtime depends mostly on the size of the closure, so deep class hierarchies and heavy use of owl:sameAs cost more than the base triple count suggests.
The reasoner runs in-process on a single core. Loading scales to billions of triples, but reasoning runs on a graph sized for your hardware. Use carve_graph to copy out the part of a large graph you need to reason over. Scale of reasoning has measured numbers.
Run materialization as a batch step, after loading and before the queries that need it, rather than inside an interactive request.