Scale of reasoning
The two halves of pgRDF scale differently:
Loading is parallel and scales to billions of triples. Reasoning is single-threaded per graph, so you reason over a graph sized for your hardware.
Right-size the graph you reason over
pgrdf.materialize runs the reasoner in-process on a single core. Forward-chaining a closure is CPU-bound, so extra cores don't speed up one materialize call.
The staged loader can load the complete 8.2-billion-triple Wikidata "truthy" dump into one PostgreSQL instance. Reasoning over all of it is a different job. Reason over the part of the data your rules need, at a size your machine can close in your batch window. carve_graph copies that part out, either every triple with a given predicate or the neighbourhood of some seed nodes, into its own graph.
Benchmark: LUBM
Two reference points show what "right-sized" means in practice. Both runs were checked against the known LUBM answer counts. They were measured on an earlier pgRDF build; treat them as orders of magnitude, not a guarantee for your hardware.
LUBM-100 on a laptop, with no tuning
The LUBM-100 benchmark (100 universities, 14 reference queries), run with no database tuning at all:
| Measured | Result |
|---|---|
| Load 13,879,970 triples (Turtle) | 3 min 29 s |
| OWL 2 RL reasoning → 22.5M facts, statistics refreshed automatically | 4 min 54 s |
| All 14 queries on the loaded graph | each ≤ 3 s |
| All 14 queries after reasoning | each ≤ 5 s |
Environment: a laptop VM (Apple silicon, 8 vCPU, 32 GiB) running PostgreSQL in Docker with the default configuration. No manual indexes, no ANALYZE, no planner hints, no pgRDF settings changed.
The LUBM ladder on one machine, up to a 112-million-quad closure
The same load → OWL 2 RL materialize → SPARQL cycle across the LUBM ladder, on one 32-vCPU / 256 GiB machine running a single PostgreSQL instance:
| LUBM-N | Asserted triples | materialize (OWL 2 RL) | Total quads after reasoning |
|---|---|---|---|
| 10 | 1.32M | 15 s | 2.13M |
| 100 | 13.9M | 4 min 37 s | 22.46M |
| 250 | 34.5M | 10 min 9 s | 55.88M |
| 500 | 69.1M | ~43 min | 111.83M |
LUBM-500 produces a materialized closure of 111.8 million quads on one machine (peak memory 146 of 256 GiB): load, reason and query in one PostgreSQL instance. materialize is the dominant cost at the top of the ladder because it is the single-threaded step. Loading and indexing use all cores.
How to right-size your reasoning
- Reason over a subgraph, not the whole dataset. Carve out the part you need with
carve_graph, materialize it, then query asserted and inferred triples together. - Use a lighter rule set. The
'rdfs'profile computes only the schema closures and costs less than full OWL 2 RL when you don't need equivalence, inverse, transitive orsameAsreasoning. - Run it as a batch step.
materializeis safe to re-run and refreshes planner statistics itself, so it fits a nightly job or the end of a load.
See also
- Mental model: the cost of forward chaining.
- Reasoning profiles: choosing a lighter rule set.
- Staged loader: the parallel loading path.
- Managing graphs: carving a subgraph.