Skip to content

content_copyCarve ​

Copy a right-sized slice of a large graph into a graph of its own, so Reason and Validate run over a graph that fits your hardware instead of the whole source.

How you run it ​

sql
pgrdf.carve_graph(src BIGINT, predicate TEXT, dst BIGINT)                          → BIGINT
pgrdf.carve_graph(src BIGINT, seeds TEXT[], dst BIGINT, max_hops INT DEFAULT 1)    → BIGINT

Both forms copy triples from src into dst and return the number copied. The source is not changed.

  • By predicate: every triple whose predicate is predicate.
  • By neighbourhood: every triple within max_hops steps of the seeds, following triples in either direction.

A worked example ​

A small catalogue: a class hierarchy, three engineers, two projects.

sql
SELECT pgrdf.add_graph('http://example.org/catalogue');   -- → 1
SELECT pgrdf.parse_turtle('
@prefix ex:   <http://example.org/> .
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .

ex:Engineer rdfs:subClassOf ex:Person .
ex:Person   rdfs:subClassOf ex:Agent .

ex:alice a ex:Engineer ; ex:name "Alice" ; ex:worksOn ex:apollo .
ex:bob   a ex:Engineer ; ex:name "Bob"   ; ex:worksOn ex:apollo .
ex:carol a ex:Engineer ; ex:name "Carol" ; ex:worksOn ex:zeus .
ex:apollo ex:partOf ex:space .
ex:zeus   ex:partOf ex:olympus .
', pgrdf.graph_id('http://example.org/catalogue'));      -- → 13

By predicate ​

Carve out the class hierarchy:

sql
SELECT pgrdf.add_graph('http://example.org/schema');
SELECT pgrdf.carve_graph(
  pgrdf.graph_id('http://example.org/catalogue'),
  'http://www.w3.org/2000/01/rdf-schema#subClassOf',
  pgrdf.graph_id('http://example.org/schema'));
-- → 2

SELECT * FROM pgrdf.export_graph(pgrdf.graph_id('http://example.org/schema'));
-- <http://example.org/Engineer> <http://www.w3.org/2000/01/rdf-schema#subClassOf> <http://example.org/Person> .
-- <http://example.org/Person> <http://www.w3.org/2000/01/rdf-schema#subClassOf> <http://example.org/Agent> .

By neighbourhood ​

Carve everything within one step of alice:

sql
SELECT pgrdf.add_graph('http://example.org/alice-1hop');
SELECT pgrdf.carve_graph(
  pgrdf.graph_id('http://example.org/catalogue'),
  ARRAY['http://example.org/alice'],
  pgrdf.graph_id('http://example.org/alice-1hop'));
-- NOTICE:  carve_graph: neighbourhood continues beyond max_hops=1 — 4 adjacent node(s)
--          were not expanded; the slice is the requested 1-hop ball, not a closed
--          component (raise max_hops to widen)
-- → 8

SELECT * FROM pgrdf.export_graph(pgrdf.graph_id('http://example.org/alice-1hop'));
-- <http://example.org/Engineer> <http://www.w3.org/2000/01/rdf-schema#subClassOf> <http://example.org/Person> .
-- <http://example.org/alice> <http://example.org/name> "Alice" .
-- <http://example.org/alice> <http://example.org/worksOn> <http://example.org/apollo> .
-- <http://example.org/alice> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Engineer> .
-- <http://example.org/apollo> <http://example.org/partOf> <http://example.org/space> .
-- <http://example.org/bob> <http://example.org/worksOn> <http://example.org/apollo> .
-- <http://example.org/bob> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Engineer> .
-- <http://example.org/carol> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Engineer> .

Read the result like this:

  • alice's own triples, plus every triple that touches a node one step away: ex:Engineer and ex:apollo.
  • The NOTICE says the slice stopped at the hop limit rather than at the edge of the data. Raise max_hops to widen it. With max_hops => 5, this slice takes all 13 triples and there is no NOTICE.
  • Busy nodes pull in a lot. ex:Engineer is one step from alice, so every engineer's rdf:type triple comes along, carol's included. In a large graph a popular class can pull in many triples at one hop.
  • Predicates are not nodes. The walk goes from node to node along triples. An axiom about a property, such as ex:manages owl:inverseOf ex:reportsTo, is not reached through ex:manages being used as a predicate. Carve such axioms by predicate as well; the carve pattern shows how.

The destination ​

  • Create it first with add_graph(iri), then pass graph_id(iri). That gives the slice a name you can query with GRAPH <iri>.

  • The id form creates a missing destination. Passing a plain id for a graph that doesn't exist creates it, named urn:pgrdf:graph:<id>:

    sql
    SELECT pgrdf.carve_graph(pgrdf.graph_id('http://example.org/catalogue'),
                             'http://example.org/partOf', 50);            -- → 2
    SELECT graph_id, iri, asserted FROM pgrdf.graph_inventory() WHERE graph_id = 50;
    --  graph_id |        iri         | asserted
    -- ----------+--------------------+----------
    --        50 | urn:pgrdf:graph:50 |        2
  • Carve appends. Carving into a graph that already holds triples adds to it, and a triple already there is stored a second time. Carving the schema again above leaves schema with 4 triples, not 2. Carve each slice into a fresh, empty graph.

  • Watch for NULL. graph_id('…') returns NULL for an IRI that doesn't exist, and carve_graph given a NULL argument returns NULL without copying anything.

  • Source and destination must differ (22023).

  • A locked destination refuses with SQLSTATE 55P03. A locked source is fine: carving only reads it.

    sql
    SELECT pgrdf.lock_graph(pgrdf.graph_id('http://example.org/schema'), 'published');
    SELECT pgrdf.carve_graph(pgrdf.graph_id('http://example.org/catalogue'),
      'http://www.w3.org/2000/01/rdf-schema#subClassOf', pgrdf.graph_id('http://example.org/schema'));
    -- ERROR:  55P03: pgrdf: graph 2 is locked (published): carve_graph (destination) refused.
    --         Unlock with pgrdf.unlock_graph(2, '<reason>').
  • Inferred triples come along. If the source has been materialized, the inferred triples inside the slice are copied too, and the slice's inventory shows materialization = unknown. Run materialize on the slice to derive them for the slice itself.

Typical use ​

Load the full source, carve the part you care about into a new graph, then work on the slice:

sql
SELECT pgrdf.add_graph('http://example.org/slice');
SELECT pgrdf.carve_graph(pgrdf.graph_id('http://example.org/catalogue'),
                         ARRAY['http://example.org/apollo'],
                         pgrdf.graph_id('http://example.org/slice'), 3);
SELECT pgrdf.materialize(pgrdf.graph_id('http://example.org/slice'));

The Ingest → Carve → Reason pattern walks through this end to end.

Where it sits in a chain ​

After Import (and optionally Seal of the source), before Reason. Once the slice is done, the source can be unloaded.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.