Skip to content

Hexastore + dictionary ​

Every IRI, blank node and literal is stored once, in a dictionary. Every triple is stored as four integer ids and indexed three ways (SPO, POS, OSP), so every triple pattern has an index to use.

What it does ​

The storage layout combines two ideas:

  1. Dictionary encoding. Each distinct term (an IRI, a blank node, or a literal with its datatype and language) is stored once and given an integer id. A triple is stored as four ids (subject, predicate, object, graph), not four strings.
  2. Three covering indexes. The quad table carries three B-tree indexes, SPO, POS and OSP, so whichever positions of a SPARQL triple pattern are bound, one index serves the lookup.

Together these keep storage compact (no repeated strings) and make triple-pattern lookups index scans. The layout holds at billion scale: the full 8.2-billion-triple Wikidata graph was loaded into a single PostgreSQL instance with the complete three-index set; see the staged loader for the measurements.

Why you'd use it ​

  • Project managers — storage grows with the number of distinct terms, not with how often they repeat. A predicate used a million times is stored once.
  • Data scientists — SPARQL patterns become integer joins over indexed columns; there is no planner tuning to do.
  • Ontologists — dense rdf:type and rdfs:subClassOf structures take a fraction of the space of a text-keyed table.

Example ​

pgrdf.sparql_sql() shows the SQL a SPARQL query becomes, which makes the encoding visible:

sql
SELECT pgrdf.add_graph('http://example.org/people');
SELECT pgrdf.parse_turtle('
@prefix ex:   <http://example.org/> .
@prefix foaf: <http://xmlns.com/foaf/0.1/> .
ex:alice a foaf:Person ; foaf:name "Alice" .
ex:bob   a foaf:Person ; foaf:name "Bob" .
', pgrdf.graph_id('http://example.org/people'));

SELECT pgrdf.sparql_sql('
  PREFIX foaf: <http://xmlns.com/foaf/0.1/>
  SELECT ?s WHERE { ?s a foaf:Person }');
--  SELECT (SELECT lexical_value FROM pgrdf._pgrdf_dictionary
--                 WHERE id = q1.subject_id) AS "s" FROM pgrdf._pgrdf_quads q1 WHERE q1.predicate_id = 4 AND q1.object_id = 8

SELECT pgrdf.get_term(4), pgrdf.get_term(8);
--  http://www.w3.org/1999/02/22-rdf-syntax-ns#type | http://xmlns.com/foaf/0.1/Person

The pattern ?s a foaf:Person has its predicate and object bound and its subject free, so it is an integer lookup on the POS index; subjects are turned back into text only for the result. pgrdf.get_term(id) returns the text for any term id. (The ids differ from database to database.)

How it works ​

  • The dictionary enforces one row per distinct term, keyed on the full term identity (kind, value, datatype, language). Very long literals are handled by keying on a fixed-size hash of the value.
  • The quad table has a column per position plus an is_inferred flag, so triples written by materialize sit beside asserted ones and can be told apart. It is partitioned by graph; see Per-graph partitions.
  • The SPO, POS and OSP indexes are declared on the partitioned table, so every graph's partition has all three.

The tables themselves are internal; use the SQL functions and graph_inventory() rather than querying them directly. For the full layout, see Storage internals.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.