Idempotence and scheduling
pgrdf.materialize(g)can be re-run as often as you like. Each run replaces the previous inferred triples. Asserted triples are never touched.
What you can rely on
| Behaviour | What you see |
|---|---|
| Re-running doesn't duplicate inferred triples | previous_inferred_dropped equals the previous run's inferred_triples_written, and the inferred count in graph_inventory() stays the same. |
| Asserted triples are never changed | asserted in graph_inventory() and the export_graph output are identical before and after. |
| A run after an edit reflects the edited graph | The reasoner reads the current asserted triples each time. |
| The call is transactional | Inside BEGIN … ROLLBACK, the rollback undoes the new inferred triples along with everything else. |
| An empty graph is fine | base_triples is 0. With 'owl-rl', inferred_triples_written is 4: the OWL axiomatic triples about owl:Thing and owl:Nothing. |
| A locked graph refuses | SQLSTATE 55P03, so a scheduled job fails loudly rather than skipping silently. |
Worked example
SELECT pgrdf.materialize(pgrdf.graph_id('urn:example:people'));
-- {..., "base_triples": 3, "inferred_triples_written": 10, "previous_inferred_dropped": 0}
SELECT pgrdf.materialize(pgrdf.graph_id('urn:example:people'));
-- {..., "base_triples": 3, "inferred_triples_written": 10, "previous_inferred_dropped": 10}
-- Same result. The previous inferred set was replaced, not added to.
-- Edit the asserted triples.
SELECT pgrdf.parse_turtle('
@prefix ex: <http://example.com/> .
ex:bob a ex:Engineer .
', pgrdf.graph_id('urn:example:people'));
SELECT pgrdf.materialize(pgrdf.graph_id('urn:example:people'));
-- {..., "base_triples": 4, "inferred_triples_written": 13, "previous_inferred_dropped": 10}Scheduling it
Because there is no "first run" versus "later run" to distinguish, you can call materialize:
- on a schedule, for example nightly from
pg_cronor an external job runner; - at the end of a load, as the last statement of the loading transaction or script;
- by hand during development, as often as you like.
Planner statistics are refreshed after each run (auto_analyzed in the result), so you don't need a separate ANALYZE.
To refresh only the graphs that need it, drive the job from graph_inventory(). Here urn:example:people had new triples loaded since its last run (stale), and urn:example:people-copy was filled by copy_graph (unknown):
SELECT iri, pgrdf.materialize(graph_id)->>'inferred_triples_written' AS written
FROM pgrdf.graph_inventory()
WHERE materialization IN ('stale', 'unknown') AND NOT locked;
-- iri | written
-- -------------------------+---------
-- urn:example:people | 16
-- urn:example:people-copy | 13Both graphs then read current.
Graphs that have never been materialized read never and are left alone by this query. Add them explicitly if you want them reasoned over. Remember the freshness limit: an edit that leaves the asserted count unchanged still reads current.
Size the graph, not just the schedule
Reasoning is single-threaded per graph, so a scheduled job should run on a graph your hardware can close within the batch window. See Scale of reasoning.
Removing inferred triples
There is no separate call to delete only the inferred triples. In practice you rarely need one:
- to bring inferred triples up to date, run
materializeagain. It replaces them; - to work with the asserted triples alone, use
export_graph, which never includes inferred triples; - to drop a graph together with its inferred triples, use
drop_graph(g). The defaultcascade => truecovers them.
Don't delete rows from pgRDF's internal tables (pgrdf._pgrdf_*). They aren't a supported interface, and editing them directly bypasses the freshness tracking that graph_inventory() reports.