storageImport
Get RDF into a graph. Import is the first verb in every chain, and the one that scales furthest: the staged loader ingested the complete 8.2-billion-triple Wikidata dump into one instance.
What it is
Import parses RDF, stores each distinct term once in a shared dictionary, and writes the triples into the graph you name. The loaders build the indexes as they go, so a graph is ready to query as soon as the call returns. There's no separate index-building step. Input can be Turtle, N-Triples, TriG or N-Quads.
Pick an entry point
| You have | Use |
|---|---|
| RDF in a string (tests, fixtures, application data) | parse_turtle, parse_trig, parse_nquads |
| A Turtle or N-Triples file on the database server | load_turtle |
| A very large N-Triples file and an empty database | load_turtle_staged_run |
| A file larger than memory | load_turtle_streaming |
Create the target graph first. add_graph returns its id, and graph_id looks it up again later:
SELECT pgrdf.add_graph('http://example.org/books'); -- → 1load_turtle — files on the server
pgrdf.load_turtle(path TEXT, graph_id BIGINT,
base_iri TEXT DEFAULT NULL,
bulk_load BOOLEAN DEFAULT false) → BIGINTReads a .ttl or .nt file from the database server's filesystem, not your client's, and returns the number of triples loaded. load_turtle_verbose takes the same arguments and returns a JSON report instead: the strategy used, triples, timings and parse_skipped (lines passed over).
SELECT pgrdf.load_turtle('/data/books.ttl', pgrdf.graph_id('http://example.org/books'));
-- NOTICE: pgrdf.load_turtle: input is Turtle (prefixed/multi-line); using the full parser.
-- For the faster staged loader, supply N-Triples (one bare-term statement per line)
-- with pgrdf preloaded
-- → 4The NOTICE is advice, not an error: Turtle goes through the full parser. See Load Turtle from disk.
Leave bulk_load off for Turtle
bulk_load => true is for N-Triples only. Given prefixed or multi-line Turtle, it can load zero triples without raising an error. parse_skipped in the verbose report counts the lines it passed over. Leave it at the default for Turtle, and check the returned count.
load_turtle_staged_run — very large N-Triples files
pgrdf.load_turtle_staged_run(path TEXT, graph_id BIGINT, n_workers INT DEFAULT 0) → JSONBThe parallel loader. It runs in four phases (stage, dictionary, resolve, index) spread across background workers. n_workers => 0 sizes the pool to the host.
SELECT pgrdf.load_turtle_staged_run('/data/books.nt', pgrdf.graph_id('http://example.org/books'));
-- → {"ok": true, "quads": 4, "job_id": 19, "triples": 4,
-- "phase_ms": {"dict": 37.316066, "index": 14.003394, "stage": 10.646233, "resolve": 9.472735},
-- "n_workers": 4, "dict_terms": 10}CALL pgrdf.load_turtle_staged(path, graph_id, n_workers) is the procedure form. It reports the same JSON as a NOTICE.
Its rules:
N-Triples only: one statement per line. Malformed lines are skipped and counted.
An empty database. The term dictionary must be empty, so run it in a database with no RDF loaded yet. Otherwise it loads nothing and says so:
json{"ok": false, "fallback": true, "reason": "dictionary already populated — staged loader requires an empty dict; caller should use the combined path", ...}Check
ok. When it isfalse, load the file withload_turtle.Not inside a transaction block. It commits after each phase, so it refuses with an error inside
BEGIN … COMMIT. Call it as a single statement.shared_preload_libraries = 'pgrdf'must be set. See Install.
See Native staged bulk loader.
load_turtle_streaming — bounded windows
pgrdf.load_turtle_streaming(path TEXT, graph_id BIGINT,
window_triples INT DEFAULT 20000000,
id_reserve_block INT DEFAULT 1000000,
base_iri TEXT DEFAULT NULL) → JSONBReads the file in windows of window_triples, for inputs larger than memory, and returns the same report as load_turtle_verbose. See Bulk ingest.
parse_turtle, parse_trig, parse_nquads — inline strings
pgrdf.parse_turtle(content TEXT, graph_id BIGINT, base_iri TEXT DEFAULT NULL) → BIGINT
pgrdf.parse_trig (content TEXT, default_graph_id BIGINT DEFAULT 0, strict BOOLEAN DEFAULT false) → JSONB
pgrdf.parse_nquads(content TEXT, default_graph_id BIGINT DEFAULT 0, strict BOOLEAN DEFAULT false) → JSONBparse_turtle ingests a Turtle string into one graph and returns the triple count. Relative IRIs such as <n1> need a base_iri. Without one, the call refuses with a parse error (No scheme found in an absolute IRI).
SELECT pgrdf.add_graph('http://example.org/notes');
SELECT pgrdf.parse_turtle('<n1> <http://example.org/text> "hello" .',
pgrdf.graph_id('http://example.org/notes'),
'http://example.org/');
-- → 1 (stored as <http://example.org/n1>)TriG and N-Quads carry their own graph names. With strict => false (the default), parse_trig and parse_nquads create any named graph that doesn't exist yet, and the report's graphs key lists the ids they wrote to:
SELECT pgrdf.parse_trig('
@prefix ex: <http://example.org/> .
ex:shelf-1 { ex:b1 ex:onShelf ex:shelf-1 . }
ex:shelf-2 { ex:b2 ex:onShelf ex:shelf-2 . }
');
SELECT iri, asserted FROM pgrdf.graph_inventory() WHERE iri LIKE '%shelf%';
-- iri | asserted
-- ----------------------------+----------
-- http://example.org/shelf-1 | 1
-- http://example.org/shelf-2 | 1With strict => true, an unknown graph name is refused (42704 parse_trig: unknown graph iri …). See Inline Turtle ingest.
Things to know
- Reloading appends. Loading the same file into the same graph a second time stores its triples again. To reload a graph from source,
clear_graphit first. - RDF-star quoted triples are rejected.
- A locked graph refuses every load with SQLSTATE
55P03. See Seal.
Where it sits in a chain
First. Every chain begins with Import. What follows depends on the chain: Seal to freeze a source before Carve, or straight into Reason, Validate or Query.
Scaling class — parallel
The staged loader spreads its phases across a pool of background workers, so more cores load faster. It is the part of the operating model that scales with the machine, up to the 8.2-billion-triple ingest.
See also
- Pillar 1 — Semantic storage — the storage layer Import writes into.
- Native staged bulk loader — the engine behind very large imports.
- Managing graphs — create, inventory, clear and drop graphs.