Hexastore + dictionary
Every IRI, blank node and literal is stored once, in a dictionary. Every triple is stored as four integer ids and indexed three ways (SPO, POS, OSP), so every triple pattern has an index to use.
What it does
The storage layout combines two ideas:
- Dictionary encoding. Each distinct term (an IRI, a blank node, or a literal with its datatype and language) is stored once and given an integer id. A triple is stored as four ids (subject, predicate, object, graph), not four strings.
- Three covering indexes. The quad table carries three B-tree indexes, SPO, POS and OSP, so whichever positions of a SPARQL triple pattern are bound, one index serves the lookup.
Together these keep storage compact (no repeated strings) and make triple-pattern lookups index scans. The layout holds at billion scale: the full 8.2-billion-triple Wikidata graph was loaded into a single PostgreSQL instance with the complete three-index set; see the staged loader for the measurements.
Why you'd use it
- Project managers — storage grows with the number of distinct terms, not with how often they repeat. A predicate used a million times is stored once.
- Data scientists — SPARQL patterns become integer joins over indexed columns; there is no planner tuning to do.
- Ontologists — dense
rdf:typeandrdfs:subClassOfstructures take a fraction of the space of a text-keyed table.
Example
pgrdf.sparql_sql() shows the SQL a SPARQL query becomes, which makes the encoding visible:
SELECT pgrdf.add_graph('http://example.org/people');
SELECT pgrdf.parse_turtle('
@prefix ex: <http://example.org/> .
@prefix foaf: <http://xmlns.com/foaf/0.1/> .
ex:alice a foaf:Person ; foaf:name "Alice" .
ex:bob a foaf:Person ; foaf:name "Bob" .
', pgrdf.graph_id('http://example.org/people'));
SELECT pgrdf.sparql_sql('
PREFIX foaf: <http://xmlns.com/foaf/0.1/>
SELECT ?s WHERE { ?s a foaf:Person }');
-- SELECT (SELECT lexical_value FROM pgrdf._pgrdf_dictionary
-- WHERE id = q1.subject_id) AS "s" FROM pgrdf._pgrdf_quads q1 WHERE q1.predicate_id = 4 AND q1.object_id = 8
SELECT pgrdf.get_term(4), pgrdf.get_term(8);
-- http://www.w3.org/1999/02/22-rdf-syntax-ns#type | http://xmlns.com/foaf/0.1/PersonThe pattern ?s a foaf:Person has its predicate and object bound and its subject free, so it is an integer lookup on the POS index; subjects are turned back into text only for the result. pgrdf.get_term(id) returns the text for any term id. (The ids differ from database to database.)
How it works
- The dictionary enforces one row per distinct term, keyed on the full term identity (kind, value, datatype, language). Very long literals are handled by keying on a fixed-size hash of the value.
- The quad table has a column per position plus an
is_inferredflag, so triples written bymaterializesit beside asserted ones and can be told apart. It is partitioned by graph; see Per-graph partitions. - The SPO, POS and OSP indexes are declared on the partitioned table, so every graph's partition has all three.
The tables themselves are internal; use the SQL functions and graph_inventory() rather than querying them directly. For the full layout, see Storage internals.
See also
- Native staged bulk loader — what writes this layout at billion scale.
- Term types — the kinds of term the dictionary holds.
- Shared-memory term cache.