Skip to content

hubPattern — Ingest → Carve → Reason ​

The chain for when the source graph is larger than one session can reason over. Load all of it, which scales, carve the slice you actually need into its own graph, and reason over the slice.

Green is parallel, amber runs in one session over one graph, grey is a single graph operation, and blue is the output.

When to use it ​

The source is bigger than one session can close over, up to the 8.2-billion-triple extreme. You still load all of it, because loading scales, but you reason over a slice sized to your hardware.

A worked scenario — reason over one project's team ​

The walkthrough uses a 15-triple source so every output fits on the page. The calls are the same at any size.

Step 1 — load the source ​

sql
SELECT pgrdf.add_graph('http://example.org/source');   -- → 1
SELECT pgrdf.parse_turtle('
@prefix ex:   <http://example.org/> .
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .
@prefix owl:  <http://www.w3.org/2002/07/owl#> .

ex:Engineer rdfs:subClassOf ex:Employee .
ex:Employee rdfs:subClassOf ex:Person .
ex:manages  owl:inverseOf   ex:reportsTo .

ex:alice a ex:Engineer ; ex:name "Alice" ; ex:worksOn ex:apollo .
ex:bob   a ex:Engineer ; ex:name "Bob"   ; ex:worksOn ex:apollo ; ex:manages ex:alice .
ex:carol a ex:Engineer ; ex:name "Carol" ; ex:worksOn ex:zeus .
ex:apollo ex:partOf ex:space .
ex:zeus   ex:partOf ex:olympus .
', pgrdf.graph_id('http://example.org/source'));      -- → 15

At scale

A large source arrives as an N-Triples file and goes into an empty database through the parallel staged loader:

sql
SELECT pgrdf.load_turtle_staged_run('/data/source.nt', pgrdf.graph_id('http://example.org/source'));
-- → {"ok": true, "triples": ..., "quads": ..., "phase_ms": {...}, ...}

See Import for its rules.

Step 2 — seal the source ​

Lock the source so it can't change while you carve from it, and record its fingerprint so the slice can say which version it came from:

sql
SELECT pgrdf.lock_graph(pgrdf.graph_id('http://example.org/source'), 'source for carving');
SELECT pgrdf.graph_digest(pgrdf.graph_id('http://example.org/source'));
-- → 7a1cbeb616157bbbaf18e54f41faea8b1e46e0f6125c6c81b1660541482ff865

Carving only reads the source, so it works on a locked graph.

Step 3 — carve the slice ​

Carve everything within three steps of the apollo project into a new, empty graph:

sql
SELECT pgrdf.add_graph('http://example.org/apollo-team');   -- → 2
SELECT pgrdf.carve_graph(
  pgrdf.graph_id('http://example.org/source'),
  ARRAY['http://example.org/apollo'],
  pgrdf.graph_id('http://example.org/apollo-team'),
  3);
-- NOTICE:  carve_graph: neighbourhood continues beyond max_hops=3 — 3 adjacent node(s)
--          were not expanded; the slice is the requested 3-hop ball, not a closed
--          component (raise max_hops to widen)
-- → 13

The NOTICE says the source goes on beyond three hops. That's expected here: you asked for a bounded slice. Check that the slice reaches as far as your reasoning needs. At max_hops => 2, this slice would stop before ex:Employee rdfs:subClassOf ex:Person, and alice would never be inferred to be a Person.

The neighbourhood walk follows nodes, not predicates, so the property axiom ex:manages owl:inverseOf ex:reportsTo isn't in the slice. Carve it by predicate into the same graph:

sql
SELECT pgrdf.carve_graph(
  pgrdf.graph_id('http://example.org/source'),
  'http://www.w3.org/2002/07/owl#inverseOf',
  pgrdf.graph_id('http://example.org/apollo-team'));
-- → 1

Carve appends, and a triple already in the destination would be stored twice. The inverse axiom wasn't in the slice yet, so nothing is doubled here.

Step 4 — reason over the slice ​

sql
SELECT pgrdf.materialize(pgrdf.graph_id('http://example.org/apollo-team'));
-- → {"profile": "owl-rl", "base_triples": 14, "inferred_triples_written": 19,
--    "previous_inferred_dropped": 0, "reasoner_errors": [], ...}

Reasoning runs over 14 triples, not the whole source.

Step 5 — query the slice ​

Scope the queries to the slice with GRAPH:

sql
SELECT * FROM pgrdf.sparql(
  'PREFIX ex: <http://example.org/>
   SELECT ?type WHERE {
     GRAPH <http://example.org/apollo-team> { ex:alice a ?type }
   } ORDER BY ?type');
json
{"type": "http://example.org/Employee"}
{"type": "http://example.org/Engineer"}
{"type": "http://example.org/Person"}
{"type": "http://www.w3.org/2002/07/owl#Thing"}
sql
SELECT * FROM pgrdf.sparql(
  'PREFIX ex: <http://example.org/>
   SELECT ?who ?boss WHERE {
     GRAPH <http://example.org/apollo-team> { ?who ex:reportsTo ?boss }
   }');
json
{"who": "http://example.org/alice", "boss": "http://example.org/bob"}

Alice's two superclasses and her reportsTo link were all derived over the slice alone. The inventory shows the split:

sql
SELECT iri, asserted, inferred, locked, materialization
  FROM pgrdf.graph_inventory() WHERE graph_id > 0;
--               iri               | asserted | inferred | locked | materialization
-- --------------------------------+----------+----------+--------+-----------------
--  http://example.org/source      |       15 |        0 | t      | never
--  http://example.org/apollo-team |       14 |       19 | f      | current

Add a Validate step on the slice here if you need a conformance gate.

Step 6 — unload the source ​

Once you have the slices you need, the source can leave the database. Package it first if you might want it back (see Unload), then unlock and drop it:

sql
SELECT pgrdf.unlock_graph(pgrdf.graph_id('http://example.org/source'), 'slice done');
SELECT pgrdf.drop_graph('http://example.org/source');
-- → 15

The slice is untouched:

sql
SELECT iri, asserted, inferred, locked, materialization
  FROM pgrdf.graph_inventory() WHERE graph_id > 0;
--               iri               | asserted | inferred | locked | materialization
-- --------------------------------+----------+----------+--------+-----------------
--  http://example.org/apollo-team |       14 |       19 | f      | current

The chain at a glance ​

StepVerbCall
1Importparse_turtle / load_turtle / load_turtle_staged_run
2Seallock_graph, graph_digest
3Carvecarve_graph by neighbourhood, then by predicate for property axioms
4Reasonmaterialize on the slice
5Querysparql with GRAPH <slice>
6Unloadexport_graph + graph_manifest, then unlock_graph + drop_graph

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.