Skip to content

query_statsScale & benchmarks ​

Two measured benchmarks: one for loading, one for reasoning.

Loading: the complete Wikidata dump ​

The staged loader loaded the complete Wikidata truthy N-Triples dump into a single PostgreSQL instance on a 128-vCPU cloud VM with 1 TiB of RAM:

MeasuredResult
Triples loaded8,199,708,346 (0 dropped)
Wall-clock time4 h 53 m
Average rate~466 K triples/s
Distinct terms in the dictionary1,801,847,593
On-disk size~2.0 TB: heap 729 GB + indexes 1,448 GB
IndexesSPO, POS and OSP

The load was checked for completeness: every triple in the dump became one stored quad. Literals are deduplicated on their full value, datatype and language, so "Berlin" in its 268 language-tagged forms is 268 distinct terms, never merged into one.

Time per phase:

PhaseTimeWhat it does
Stage14 minParse the dump into a staging table, in parallel.
Dictionary1 h 51 mDeduplicate the 1.8 billion distinct terms and assign ids.
Resolve2 h 00 mTurn staged triples into encoded quads.
Index32 minBuild the SPO, POS and OSP indexes, concurrently.

The same full load also completed on a 64-vCPU VM with 503 GiB of RAM and a 3.4 TB disk, in about 10.3 hours (~221 K triples/s), with no tuning.

Running the staged loader ​

The staged loader takes N-Triples, needs pgrdf in shared_preload_libraries, runs outside a transaction block, and loads into an empty database. With n_workers => 0 (the default) it sizes its worker pool to the machine. The file path is on the database server:

sql
SELECT pgrdf.load_turtle_staged_run('/data/people.nt', 0, n_workers => 0);
json
{"ok": true, "quads": 3, "job_id": 17, "triples": 3, "phase_ms": {"dict": 33.422963, "index": 7.4340329999999994, "stage": 8.220616999999999, "resolve": 7.811575}, "n_workers": 4, "dict_terms": 7}

The same call that loads a three-line file loads the Wikidata dump. The staged loader page covers its options.

Reasoning: LUBM-500 ​

LUBM is the standard university-domain benchmark for reasoners. On a single 32-vCPU machine with 256 GiB of RAM, pgRDF loads, reasons over and queries the LUBM ladder in one PostgreSQL instance:

LUBMBase triplesOWL 2 RL materializeQuads after reasoning
10013.9 M4 m 37 s22.46 M
25034.5 M10 m 9 s55.88 M
50069.1 M~43 m111.83 M

Loading and reasoning scale differently

The Wikidata figure is loading, not reasoning: truthy statements are already direct claims, so there is nothing to infer. Loading spreads over many cores. materialize runs in one backend, so reason over a graph sized to the machine, carving it out of a larger one when needed. See Scale of reasoning and Processes & flows.

See also ​

  • bolt Staged bulk loader: the loader behind the Wikidata numbers.
  • psychology Scale of reasoning: the LUBM runs in detail, and how to size reasoning.
  • account_tree Graphs: carving a working graph out of a large one.

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.