Skip to content

storagePillar 1 — Semantic storage ​

RDF storage in PostgreSQL. Turtle, N-Triples, TriG or N-Quads go in; dictionary-encoded, hexastore-indexed quads come out, one partition per named graph, queryable with SPARQL or SQL from any Postgres client.

Billion-scale ingest

The staged bulk loader loaded the complete 8.2-billion-triple Wikidata "truthy" dump into a single PostgreSQL instance, dictionary-encoded with a full SPO/POS/OSP index set. It runs as a pool of background workers, commits after each phase (STAGE → DICT → RESOLVE → INDEX) and sizes itself to the host.

Features in this pillar ​

  • description Load Turtle from disk — read a .ttl / .nt file from the database server into a graph; on a preloaded server, N-Triples go to the staged loader.
  • bolt Native staged bulk loader — load_turtle_staged_run: a multi-worker, commit-per-phase loader for very large N-Triples files.
  • description Inline Turtle / TriG / N-Quads ingest — load RDF from a SQL string; parse_trig / parse_nquads for formats that name their graphs.
  • query_stats Verbose ingest statistics — a JSONB report of the loader path, triples, skipped lines and timings.
  • storage Per-graph partitions — each graph is its own partition: cheap whole-graph operations, isolated namespaces.
  • account_tree Named graphs (IRI and id) — create graphs by IRI and move between IRIs and ids.
  • account_tree Hexastore + dictionary — every term stored once, every triple indexed three ways (SPO/POS/OSP).
  • description Term types — typed literals, language tags, blank nodes, RDF collections.
  • bolt Bulk ingest — the loader family and how to choose between the standard, bulk, streaming and staged paths.
  • bolt Shared-memory term cache — a cross-session cache for terms that keep coming back.
  • build Graph lifecycle — drop_graph, clear_graph, copy_graph, move_graph as whole-graph operations.

At a glance ​

pgRDF must be listed in shared_preload_libraries (then restart the server) for the shared term cache and the staged loader. See Install.

sql
CREATE EXTENSION pgrdf;

-- Create a graph and load some Turtle into it.
SELECT pgrdf.add_graph('http://example.org/people');
SELECT pgrdf.parse_turtle('
@prefix ex:   <http://example.org/> .
@prefix foaf: <http://xmlns.com/foaf/0.1/> .

ex:alice a foaf:Person ; foaf:name "Alice" ; foaf:knows ex:bob .
ex:bob   a foaf:Person ; foaf:name "Bob" .
', pgrdf.graph_id('http://example.org/people'));
--  → 5

-- A file on the database server loads the same way:
--   SELECT pgrdf.load_turtle('/path/on/server/people.ttl',
--                            pgrdf.graph_id('http://example.org/people'));

-- Inspect
SELECT graph_id, iri, asserted FROM pgrdf.graph_inventory();
--  graph_id |            iri            | asserted
-- ----------+---------------------------+----------
--         0 | urn:pgrdf:graph:0         |        0
--         1 | http://example.org/people |        5

SELECT * FROM pgrdf.sparql('
  PREFIX foaf: <http://xmlns.com/foaf/0.1/>
  SELECT ?name WHERE { ?person a foaf:Person ; foaf:name ?name }');
--  {"name": "Bob"}
--  {"name": "Alice"}

For more on managing graphs (inventory, copying, carving, locking), see Managing graphs.

Next — Load Turtle from disk →

Further reading ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.