Skip to content

Scale of reasoning ​

The two halves of pgRDF scale differently:

Loading is parallel and scales to billions of triples. Reasoning is single-threaded per graph, so you reason over a graph sized for your hardware.

Right-size the graph you reason over ​

pgrdf.materialize runs the reasoner in-process on a single core. Forward-chaining a closure is CPU-bound, so extra cores don't speed up one materialize call.

The staged loader can load the complete 8.2-billion-triple Wikidata "truthy" dump into one PostgreSQL instance. Reasoning over all of it is a different job. Reason over the part of the data your rules need, at a size your machine can close in your batch window. carve_graph copies that part out, either every triple with a given predicate or the neighbourhood of some seed nodes, into its own graph.

Benchmark: LUBM ​

Two reference points show what "right-sized" means in practice. Both runs were checked against the known LUBM answer counts. They were measured on an earlier pgRDF build; treat them as orders of magnitude, not a guarantee for your hardware.

LUBM-100 on a laptop, with no tuning ​

The LUBM-100 benchmark (100 universities, 14 reference queries), run with no database tuning at all:

MeasuredResult
Load 13,879,970 triples (Turtle)3 min 29 s
OWL 2 RL reasoning → 22.5M facts, statistics refreshed automatically4 min 54 s
All 14 queries on the loaded grapheach ≤ 3 s
All 14 queries after reasoningeach ≤ 5 s

Environment: a laptop VM (Apple silicon, 8 vCPU, 32 GiB) running PostgreSQL in Docker with the default configuration. No manual indexes, no ANALYZE, no planner hints, no pgRDF settings changed.

The LUBM ladder on one machine, up to a 112-million-quad closure ​

The same load → OWL 2 RL materialize → SPARQL cycle across the LUBM ladder, on one 32-vCPU / 256 GiB machine running a single PostgreSQL instance:

LUBM-NAsserted triplesmaterialize (OWL 2 RL)Total quads after reasoning
101.32M15 s2.13M
10013.9M4 min 37 s22.46M
25034.5M10 min 9 s55.88M
50069.1M~43 min111.83M

LUBM-500 produces a materialized closure of 111.8 million quads on one machine (peak memory 146 of 256 GiB): load, reason and query in one PostgreSQL instance. materialize is the dominant cost at the top of the ladder because it is the single-threaded step. Loading and indexing use all cores.

How to right-size your reasoning ​

  • Reason over a subgraph, not the whole dataset. Carve out the part you need with carve_graph, materialize it, then query asserted and inferred triples together.
  • Use a lighter rule set. The 'rdfs' profile computes only the schema closures and costs less than full OWL 2 RL when you don't need equivalence, inverse, transitive or sameAs reasoning.
  • Run it as a batch step. materialize is safe to re-run and refreshes planner statistics itself, so it fits a nightly job or the end of a load.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.