storagePillar 1 — Semantic storage
RDF storage in PostgreSQL. Turtle, N-Triples, TriG or N-Quads go in; dictionary-encoded, hexastore-indexed quads come out, one partition per named graph, queryable with SPARQL or SQL from any Postgres client.
Billion-scale ingest
The staged bulk loader loaded the complete 8.2-billion-triple Wikidata "truthy" dump into a single PostgreSQL instance, dictionary-encoded with a full SPO/POS/OSP index set. It runs as a pool of background workers, commits after each phase (STAGE → DICT → RESOLVE → INDEX) and sizes itself to the host.
Features in this pillar
- description Load Turtle from disk — read a
.ttl/.ntfile from the database server into a graph; on a preloaded server, N-Triples go to the staged loader. - bolt Native staged bulk loader —
load_turtle_staged_run: a multi-worker, commit-per-phase loader for very large N-Triples files. - description Inline Turtle / TriG / N-Quads ingest — load RDF from a SQL string;
parse_trig/parse_nquadsfor formats that name their graphs. - query_stats Verbose ingest statistics — a JSONB report of the loader path, triples, skipped lines and timings.
- storage Per-graph partitions — each graph is its own partition: cheap whole-graph operations, isolated namespaces.
- account_tree Named graphs (IRI and id) — create graphs by IRI and move between IRIs and ids.
- account_tree Hexastore + dictionary — every term stored once, every triple indexed three ways (SPO/POS/OSP).
- description Term types — typed literals, language tags, blank nodes, RDF collections.
- bolt Bulk ingest — the loader family and how to choose between the standard, bulk, streaming and staged paths.
- bolt Shared-memory term cache — a cross-session cache for terms that keep coming back.
- build Graph lifecycle —
drop_graph,clear_graph,copy_graph,move_graphas whole-graph operations.
At a glance
pgRDF must be listed in shared_preload_libraries (then restart the server) for the shared term cache and the staged loader. See Install.
CREATE EXTENSION pgrdf;
-- Create a graph and load some Turtle into it.
SELECT pgrdf.add_graph('http://example.org/people');
SELECT pgrdf.parse_turtle('
@prefix ex: <http://example.org/> .
@prefix foaf: <http://xmlns.com/foaf/0.1/> .
ex:alice a foaf:Person ; foaf:name "Alice" ; foaf:knows ex:bob .
ex:bob a foaf:Person ; foaf:name "Bob" .
', pgrdf.graph_id('http://example.org/people'));
-- → 5
-- A file on the database server loads the same way:
-- SELECT pgrdf.load_turtle('/path/on/server/people.ttl',
-- pgrdf.graph_id('http://example.org/people'));
-- Inspect
SELECT graph_id, iri, asserted FROM pgrdf.graph_inventory();
-- graph_id | iri | asserted
-- ----------+---------------------------+----------
-- 0 | urn:pgrdf:graph:0 | 0
-- 1 | http://example.org/people | 5
SELECT * FROM pgrdf.sparql('
PREFIX foaf: <http://xmlns.com/foaf/0.1/>
SELECT ?name WHERE { ?person a foaf:Person ; foaf:name ?name }');
-- {"name": "Bob"}
-- {"name": "Alice"}For more on managing graphs (inventory, copying, carving, locking), see Managing graphs.
Next — Load Turtle from disk →
Further reading
- info The RDF 1.1 Primer — RDF foundations.
- description RDF 1.1 Turtle — the format
parse_turtleandload_turtleread. - code The partitioning chapter of the PostgreSQL manual.