Bulk ingest
pgRDF has several ingest paths, each suited to a different size of load: from an ontology in a string up to a source graph far larger than memory.
The loader family
| Path | When to use | What it does |
|---|---|---|
parse_turtle / load_turtle (standard) | Most loads: ontologies, fixtures, application data, any later load into a populated database | Parses in your session, resolves terms through the shared term cache, inserts in batches. Any format; works inside a transaction. |
load_turtle(…, bulk_load => true) | A first, large N-Triples load into an empty database, where the file fits in memory | Parses in parallel, builds the dictionary in memory, rebuilds indexes after the load. |
load_turtle_streaming | A first N-Triples load of a file larger than memory | Reads the file in fixed-size windows; memory stays at one window plus the dictionary. |
load_turtle_staged_run | The largest loads, into billions of triples | A multi-worker pipeline that commits after each phase. Loaded the full 8.2-billion-triple Wikidata graph. |
The three fast paths share two rules:
- N-Triples only. They read one triple per line. Malformed lines are skipped rather than stopping the load (the bulk and streaming paths count them in
parse_skipped). On Turtle with prefixes, that means every line is skipped and nothing loads, with no error. Use the standard path for Turtle. - An empty database. They are built for the first load. The bulk and streaming paths fall back to the standard parser when terms are already loaded (the report's
paththen readscombined); the staged loader declines and reports"fallback": true.
On a preloaded server, load_turtle hands N-Triples files to the staged loader by itself; see Load Turtle from disk.
Standard path
parse_turtle and load_turtle parse the input, look each term up (first in the shared cache, then in the dictionary table), and insert the triples in batches through a prepared statement. It is the only path that reads full Turtle, and the one to use for every load after the first.
Parallel bulk path — bulk_load => true
On an empty database, parses the N-Triples file across all cores, assigns dictionary ids in memory and bulk-inserts. For loads above pgrdf.bulk_defer_index_min triples (default 100000) it also drops the indexes first and rebuilds them once at the end (defer_index in the report).
Streaming loader — load_turtle_streaming
pgrdf.load_turtle_streaming(path TEXT, graph_id BIGINT,
window_triples INT DEFAULT 20000000,
id_reserve_block INT DEFAULT 1000000,
base_iri TEXT DEFAULT NULL) → JSONBReads the file one window_triples-sized window at a time (parse, resolve terms, write) and never holds the whole file in memory. Peak memory is one window plus the dictionary, regardless of file size. Indexes are rebuilt once after the last window. Returns the same report as the verbose loaders, with pathstreaming and a windows count.
Native staged loader — load_turtle_staged_run
For the largest loads, the native staged bulk loader runs a pool of background workers through four phases (STAGE → DICT → RESOLVE → INDEX), each committed before the next. It loaded the full Wikidata "truthy" graph (8,199,708,346 triples, 1,801,847,593 distinct terms) into one PostgreSQL instance. It needs pgRDF in shared_preload_libraries and can't run inside a transaction block.
Why you'd use it
- Project managers — one ingest surface from small to very large; no separate ETL system.
- Data scientists — small loads from a notebook are one call; a dump of billions of triples loads through the same family of functions.
- Ontologists — refresh a vocabulary in a CI pipeline with one SQL statement.
Example — check which path ran
With the five-line /tmp/people.nt from the staged loader example, in a fresh database (timings removed from the output):
SELECT pgrdf.add_graph('http://example.org/people');
SELECT pgrdf.load_turtle_verbose('/tmp/people.nt',
pgrdf.graph_id('http://example.org/people'),
bulk_load => true)
- ARRAY['parse_ms','dict_ms','insert_ms','resolve_ms','index_ms','elapsed_ms'];
-- {"path": "parallel_bulk", "triples": 5, "windows": 0, "dict_terms": 0, "defer_index": false,
-- "quad_batches": 1, "dict_db_calls": 0, "parse_skipped": 0, "dict_cache_hits": 0,
-- "shmem_cache_hits": 0}path names the loader that ran, triples how many went in, and parse_skipped how many lines were skipped. triples: 0 with a non-zero parse_skipped means the file wasn't N-Triples.
See also
- Native staged bulk loader — the billion-scale path.
- Verbose ingest statistics — every key in the report.
- Shared-memory term cache.