query_statsScale & benchmarks
Two measured benchmarks: one for loading, one for reasoning.
Loading: the complete Wikidata dump
The staged loader loaded the complete Wikidata truthy N-Triples dump into a single PostgreSQL instance on a 128-vCPU cloud VM with 1 TiB of RAM:
| Measured | Result |
|---|---|
| Triples loaded | 8,199,708,346 (0 dropped) |
| Wall-clock time | 4 h 53 m |
| Average rate | ~466 K triples/s |
| Distinct terms in the dictionary | 1,801,847,593 |
| On-disk size | ~2.0 TB: heap 729 GB + indexes 1,448 GB |
| Indexes | SPO, POS and OSP |
The load was checked for completeness: every triple in the dump became one stored quad. Literals are deduplicated on their full value, datatype and language, so "Berlin" in its 268 language-tagged forms is 268 distinct terms, never merged into one.
Time per phase:
| Phase | Time | What it does |
|---|---|---|
| Stage | 14 min | Parse the dump into a staging table, in parallel. |
| Dictionary | 1 h 51 m | Deduplicate the 1.8 billion distinct terms and assign ids. |
| Resolve | 2 h 00 m | Turn staged triples into encoded quads. |
| Index | 32 min | Build the SPO, POS and OSP indexes, concurrently. |
The same full load also completed on a 64-vCPU VM with 503 GiB of RAM and a 3.4 TB disk, in about 10.3 hours (~221 K triples/s), with no tuning.
Running the staged loader
The staged loader takes N-Triples, needs pgrdf in shared_preload_libraries, runs outside a transaction block, and loads into an empty database. With n_workers => 0 (the default) it sizes its worker pool to the machine. The file path is on the database server:
SELECT pgrdf.load_turtle_staged_run('/data/people.nt', 0, n_workers => 0);{"ok": true, "quads": 3, "job_id": 17, "triples": 3, "phase_ms": {"dict": 33.422963, "index": 7.4340329999999994, "stage": 8.220616999999999, "resolve": 7.811575}, "n_workers": 4, "dict_terms": 7}The same call that loads a three-line file loads the Wikidata dump. The staged loader page covers its options.
Reasoning: LUBM-500
LUBM is the standard university-domain benchmark for reasoners. On a single 32-vCPU machine with 256 GiB of RAM, pgRDF loads, reasons over and queries the LUBM ladder in one PostgreSQL instance:
| LUBM | Base triples | OWL 2 RL materialize | Quads after reasoning |
|---|---|---|---|
| 100 | 13.9 M | 4 m 37 s | 22.46 M |
| 250 | 34.5 M | 10 m 9 s | 55.88 M |
| 500 | 69.1 M | ~43 m | 111.83 M |
Loading and reasoning scale differently
The Wikidata figure is loading, not reasoning: truthy statements are already direct claims, so there is nothing to infer. Loading spreads over many cores. materialize runs in one backend, so reason over a graph sized to the machine, carving it out of a larger one when needed. See Scale of reasoning and Processes & flows.
See also
- bolt Staged bulk loader: the loader behind the Wikidata numbers.
- psychology Scale of reasoning: the LUBM runs in detail, and how to size reasoning.
- account_tree Graphs: carving a working graph out of a large one.