Skip to content

storageImport ​

Get RDF into a graph. Import is the first verb in every chain, and the one that scales furthest: the staged loader ingested the complete 8.2-billion-triple Wikidata dump into one instance.

What it is ​

Import parses RDF, stores each distinct term once in a shared dictionary, and writes the triples into the graph you name. The loaders build the indexes as they go, so a graph is ready to query as soon as the call returns. There's no separate index-building step. Input can be Turtle, N-Triples, TriG or N-Quads.

Pick an entry point ​

You haveUse
RDF in a string (tests, fixtures, application data)parse_turtle, parse_trig, parse_nquads
A Turtle or N-Triples file on the database serverload_turtle
A very large N-Triples file and an empty databaseload_turtle_staged_run
A file larger than memoryload_turtle_streaming

Create the target graph first. add_graph returns its id, and graph_id looks it up again later:

sql
SELECT pgrdf.add_graph('http://example.org/books');   -- → 1

load_turtle — files on the server ​

sql
pgrdf.load_turtle(path TEXT, graph_id BIGINT,
                  base_iri TEXT DEFAULT NULL,
                  bulk_load BOOLEAN DEFAULT false) → BIGINT

Reads a .ttl or .nt file from the database server's filesystem, not your client's, and returns the number of triples loaded. load_turtle_verbose takes the same arguments and returns a JSON report instead: the strategy used, triples, timings and parse_skipped (lines passed over).

sql
SELECT pgrdf.load_turtle('/data/books.ttl', pgrdf.graph_id('http://example.org/books'));
-- NOTICE:  pgrdf.load_turtle: input is Turtle (prefixed/multi-line); using the full parser.
--          For the faster staged loader, supply N-Triples (one bare-term statement per line)
--          with pgrdf preloaded
-- → 4

The NOTICE is advice, not an error: Turtle goes through the full parser. See Load Turtle from disk.

Leave bulk_load off for Turtle

bulk_load => true is for N-Triples only. Given prefixed or multi-line Turtle, it can load zero triples without raising an error. parse_skipped in the verbose report counts the lines it passed over. Leave it at the default for Turtle, and check the returned count.

load_turtle_staged_run — very large N-Triples files ​

sql
pgrdf.load_turtle_staged_run(path TEXT, graph_id BIGINT, n_workers INT DEFAULT 0) → JSONB

The parallel loader. It runs in four phases (stage, dictionary, resolve, index) spread across background workers. n_workers => 0 sizes the pool to the host.

sql
SELECT pgrdf.load_turtle_staged_run('/data/books.nt', pgrdf.graph_id('http://example.org/books'));
-- → {"ok": true, "quads": 4, "job_id": 19, "triples": 4,
--    "phase_ms": {"dict": 37.316066, "index": 14.003394, "stage": 10.646233, "resolve": 9.472735},
--    "n_workers": 4, "dict_terms": 10}

CALL pgrdf.load_turtle_staged(path, graph_id, n_workers) is the procedure form. It reports the same JSON as a NOTICE.

Its rules:

  • N-Triples only: one statement per line. Malformed lines are skipped and counted.

  • An empty database. The term dictionary must be empty, so run it in a database with no RDF loaded yet. Otherwise it loads nothing and says so:

    json
    {"ok": false, "fallback": true, "reason": "dictionary already populated — staged loader requires an empty dict; caller should use the combined path", ...}

    Check ok. When it is false, load the file with load_turtle.

  • Not inside a transaction block. It commits after each phase, so it refuses with an error inside BEGIN … COMMIT. Call it as a single statement.

  • shared_preload_libraries = 'pgrdf' must be set. See Install.

See Native staged bulk loader.

load_turtle_streaming — bounded windows ​

sql
pgrdf.load_turtle_streaming(path TEXT, graph_id BIGINT,
                            window_triples INT DEFAULT 20000000,
                            id_reserve_block INT DEFAULT 1000000,
                            base_iri TEXT DEFAULT NULL) → JSONB

Reads the file in windows of window_triples, for inputs larger than memory, and returns the same report as load_turtle_verbose. See Bulk ingest.

parse_turtle, parse_trig, parse_nquads — inline strings ​

sql
pgrdf.parse_turtle(content TEXT, graph_id BIGINT, base_iri TEXT DEFAULT NULL) → BIGINT
pgrdf.parse_trig  (content TEXT, default_graph_id BIGINT DEFAULT 0, strict BOOLEAN DEFAULT false) → JSONB
pgrdf.parse_nquads(content TEXT, default_graph_id BIGINT DEFAULT 0, strict BOOLEAN DEFAULT false) → JSONB

parse_turtle ingests a Turtle string into one graph and returns the triple count. Relative IRIs such as <n1> need a base_iri. Without one, the call refuses with a parse error (No scheme found in an absolute IRI).

sql
SELECT pgrdf.add_graph('http://example.org/notes');
SELECT pgrdf.parse_turtle('<n1> <http://example.org/text> "hello" .',
                          pgrdf.graph_id('http://example.org/notes'),
                          'http://example.org/');
-- → 1   (stored as <http://example.org/n1>)

TriG and N-Quads carry their own graph names. With strict => false (the default), parse_trig and parse_nquads create any named graph that doesn't exist yet, and the report's graphs key lists the ids they wrote to:

sql
SELECT pgrdf.parse_trig('
@prefix ex: <http://example.org/> .
ex:shelf-1 { ex:b1 ex:onShelf ex:shelf-1 . }
ex:shelf-2 { ex:b2 ex:onShelf ex:shelf-2 . }
');

SELECT iri, asserted FROM pgrdf.graph_inventory() WHERE iri LIKE '%shelf%';
--             iri             | asserted
-- ----------------------------+----------
--  http://example.org/shelf-1 |        1
--  http://example.org/shelf-2 |        1

With strict => true, an unknown graph name is refused (42704 parse_trig: unknown graph iri …). See Inline Turtle ingest.

Things to know ​

  • Reloading appends. Loading the same file into the same graph a second time stores its triples again. To reload a graph from source, clear_graph it first.
  • RDF-star quoted triples are rejected.
  • A locked graph refuses every load with SQLSTATE 55P03. See Seal.

Where it sits in a chain ​

First. Every chain begins with Import. What follows depends on the chain: Seal to freeze a source before Carve, or straight into Reason, Validate or Query.

Scaling class — parallel

The staged loader spreads its phases across a pool of background workers, so more cores load faster. It is the part of the operating model that scales with the machine, up to the 8.2-billion-triple ingest.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.