Skip to content

Bulk ingest ​

pgRDF has several ingest paths, each suited to a different size of load: from an ontology in a string up to a source graph far larger than memory.

The loader family ​

PathWhen to useWhat it does
parse_turtle / load_turtle (standard)Most loads: ontologies, fixtures, application data, any later load into a populated databaseParses in your session, resolves terms through the shared term cache, inserts in batches. Any format; works inside a transaction.
load_turtle(…, bulk_load => true)A first, large N-Triples load into an empty database, where the file fits in memoryParses in parallel, builds the dictionary in memory, rebuilds indexes after the load.
load_turtle_streamingA first N-Triples load of a file larger than memoryReads the file in fixed-size windows; memory stays at one window plus the dictionary.
load_turtle_staged_runThe largest loads, into billions of triplesA multi-worker pipeline that commits after each phase. Loaded the full 8.2-billion-triple Wikidata graph.

The three fast paths share two rules:

  • N-Triples only. They read one triple per line. Malformed lines are skipped rather than stopping the load (the bulk and streaming paths count them in parse_skipped). On Turtle with prefixes, that means every line is skipped and nothing loads, with no error. Use the standard path for Turtle.
  • An empty database. They are built for the first load. The bulk and streaming paths fall back to the standard parser when terms are already loaded (the report's path then reads combined); the staged loader declines and reports "fallback": true.

On a preloaded server, load_turtle hands N-Triples files to the staged loader by itself; see Load Turtle from disk.

Standard path ​

parse_turtle and load_turtle parse the input, look each term up (first in the shared cache, then in the dictionary table), and insert the triples in batches through a prepared statement. It is the only path that reads full Turtle, and the one to use for every load after the first.

Parallel bulk path — bulk_load => true ​

On an empty database, parses the N-Triples file across all cores, assigns dictionary ids in memory and bulk-inserts. For loads above pgrdf.bulk_defer_index_min triples (default 100000) it also drops the indexes first and rebuilds them once at the end (defer_index in the report).

Streaming loader — load_turtle_streaming ​

pgrdf.load_turtle_streaming(path TEXT, graph_id BIGINT,
  window_triples INT DEFAULT 20000000,
  id_reserve_block INT DEFAULT 1000000,
  base_iri TEXT DEFAULT NULL) → JSONB

Reads the file one window_triples-sized window at a time (parse, resolve terms, write) and never holds the whole file in memory. Peak memory is one window plus the dictionary, regardless of file size. Indexes are rebuilt once after the last window. Returns the same report as the verbose loaders, with pathstreaming and a windows count.

Native staged loader — load_turtle_staged_run ​

For the largest loads, the native staged bulk loader runs a pool of background workers through four phases (STAGE → DICT → RESOLVE → INDEX), each committed before the next. It loaded the full Wikidata "truthy" graph (8,199,708,346 triples, 1,801,847,593 distinct terms) into one PostgreSQL instance. It needs pgRDF in shared_preload_libraries and can't run inside a transaction block.

Why you'd use it ​

  • Project managers — one ingest surface from small to very large; no separate ETL system.
  • Data scientists — small loads from a notebook are one call; a dump of billions of triples loads through the same family of functions.
  • Ontologists — refresh a vocabulary in a CI pipeline with one SQL statement.

Example — check which path ran ​

With the five-line /tmp/people.nt from the staged loader example, in a fresh database (timings removed from the output):

sql
SELECT pgrdf.add_graph('http://example.org/people');

SELECT pgrdf.load_turtle_verbose('/tmp/people.nt',
                                 pgrdf.graph_id('http://example.org/people'),
                                 bulk_load => true)
       - ARRAY['parse_ms','dict_ms','insert_ms','resolve_ms','index_ms','elapsed_ms'];
--  {"path": "parallel_bulk", "triples": 5, "windows": 0, "dict_terms": 0, "defer_index": false,
--   "quad_batches": 1, "dict_db_calls": 0, "parse_skipped": 0, "dict_cache_hits": 0,
--   "shmem_cache_hits": 0}

path names the loader that ran, triples how many went in, and parse_skipped how many lines were skipped. triples: 0 with a non-zero parse_skipped means the file wasn't N-Triples.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.