Skip to content

Per-graph partitions ​

Each graph is its own PostgreSQL partition of the quad table. Creating one is one call; clearing or dropping a whole graph works on that partition instead of scanning rows.

What it does ​

pgrdf.add_graph(iri TEXT) → BIGINT                    -- create the graph and its partition; returns its id
pgrdf.count_quads(graph_id BIGINT DEFAULT 0) → BIGINT -- asserted + inferred triples in one graph
pgrdf.graph_inventory() → TABLE (…)                   -- every graph, with its counts

pgRDF keeps all triples in one quad table, LIST-partitioned by graph. add_graph(iri) registers the graph and creates its partition; calling it again returns the existing graph's id.

What the partitioning gives you:

  • Whole-graph operations are cheap.clear_graph and drop_graph work on one partition, so their cost doesn't grow with the graph's size.
  • Graph-scoped queries read only their graph. A SPARQL GRAPH <iri> { … } pattern becomes a condition on the graph id, and PostgreSQL skips every other partition.
  • Every partition has the full index set (SPO, POS, OSP); see Hexastore + dictionary.
  • Ordinary PostgreSQL tooling applies. Partitions are normal tables, so pg_dump includes every graph's triples.

Create the graph before loading into it

Load into pgrdf.graph_id(iri) of a graph you created with add_graph(iri). Triples loaded into a raw number that isn't a registered graph are stored in a catch-all partition that graph_inventory(), GRAPH queries and the lifecycle functions don't see.

Why you'd use it ​

  • Project managers — multi-tenant and multi-ontology workloads get cheap, isolated namespaces. Removing a tenant is one partition drop.
  • Data scientists — keep a live graph and a frozen snapshot side by side and query either, or both, in one session.
  • Ontologists — keep each loaded vocabulary in its own graph for clean composition and lifecycle.

Example ​

sql
SELECT pgrdf.add_graph('http://example.org/tenant/acme');     -- → 1
SELECT pgrdf.add_graph('http://example.org/tenant/globex');   -- → 2

SELECT pgrdf.parse_turtle('
@prefix ex: <http://example.org/> .
ex:acme ex:plan "gold" .
', pgrdf.graph_id('http://example.org/tenant/acme'));
-- → 1
SELECT pgrdf.parse_turtle('
@prefix ex: <http://example.org/> .
ex:globex ex:plan "silver" ; ex:seats 40 .
', pgrdf.graph_id('http://example.org/tenant/globex'));
-- → 2

SELECT pgrdf.count_quads(pgrdf.graph_id('http://example.org/tenant/globex'));
-- → 2

SELECT graph_id, iri, asserted FROM pgrdf.graph_inventory();
--  graph_id |               iri                | asserted
-- ----------+----------------------------------+----------
--         0 | urn:pgrdf:graph:0                |        0
--         1 | http://example.org/tenant/acme   |        1
--         2 | http://example.org/tenant/globex |        2

-- Offboard a tenant: one partition goes.
SELECT pgrdf.drop_graph('http://example.org/tenant/globex');
-- → 2

SELECT * FROM pgrdf.orphan_partitions();
-- (0 rows)

pgrdf.orphan_partitions() lists partitions that no longer belong to any graph. It should return no rows; see the inventory.

How it works ​

The quad table is declared PARTITION BY LIST (graph_id), and add_graph creates the partition for a new graph. The three covering indexes are declared on the parent table, so each partition carries them. The graph list itself is a private table: read it through graph_inventory(). For the full layout, see Storage internals.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.