Skip to content

Verbose ingest statistics ​

A JSONB report from an ingest run: which loader path ran, how many triples went in, how many lines were skipped, how much dictionary work it took, and where the time went.

What it does ​

pgrdf.parse_turtle_verbose(content TEXT, graph_id BIGINT, base_iri TEXT DEFAULT NULL) → JSONB
pgrdf.load_turtle_verbose(path TEXT, graph_id BIGINT, base_iri TEXT DEFAULT NULL,
                          bulk_load BOOLEAN DEFAULT FALSE) → JSONB

These load exactly like parse_turtle and load_turtle, but return a report instead of a bare count. parse_trig, parse_nquads and load_turtle_streaming return the same report; the quad formats add a "graphs" array.

Why you'd use it ​

  • Project managers — measure ingest cost on your own data and plan capacity from observed numbers.
  • Data scientists — check that a fast path actually ran before committing to a large load.
  • Operators — see whether time goes on parsing, dictionary lookups or inserting.

Example ​

sql
SELECT pgrdf.add_graph('http://example.org/people');

SELECT jsonb_pretty(pgrdf.parse_turtle_verbose('
@prefix ex:   <http://example.org/> .
@prefix foaf: <http://xmlns.com/foaf/0.1/> .
ex:alice a foaf:Person ; foaf:name "Alice" ; foaf:knows ex:bob .
ex:bob   a foaf:Person ; foaf:name "Bob" .
', pgrdf.graph_id('http://example.org/people')));
json
{
    "path": "combined",
    "triples": 5,
    "parse_skipped": 0,
    "quad_batches": 1,
    "dict_db_calls": 9,
    "dict_cache_hits": 8,
    "shmem_cache_hits": 0,
    "dict_terms": 0,
    "windows": 0,
    "defer_index": false,
    "parse_ms": 0.04,
    "dict_ms": 0.70,
    "insert_ms": 0.19,
    "resolve_ms": 0.0,
    "index_ms": 0.0,
    "elapsed_ms": 0.94
}

(Keys reordered for reading, timings rounded; timings vary by machine.)

Reading the report ​

KeyMeaning
pathWhich loader ran: combined (the standard parser), parallel_bulk (bulk_load => true) or streaming (load_turtle_streaming).
triplesTriples loaded.
parse_skippedLines the N-Triples fast paths skipped as unreadable. Always check this after a bulk_load or streaming run.
quad_batchesNumber of batched inserts used to write the triples.
dict_db_callsTerm lookups that went to the dictionary table (mostly new terms).
dict_cache_hitsTerm lookups answered from memory without a table lookup.
shmem_cache_hitsTerm lookups answered by the shared-memory term cache.
dict_terms, windows, defer_indexWork done by the bulk and streaming paths: terms added in memory, windows processed, whether indexes were rebuilt after the load.
parse_ms, dict_ms, insert_ms, resolve_ms, index_msTime per phase, in milliseconds.
elapsed_msTotal time inside the call, in milliseconds.

Two readings worth knowing:

  • triples: 0 with parse_skipped > 0 after bulk_load => true means the input wasn't N-Triples, so nothing loaded. Load Turtle without bulk_load.
  • dict_ms close to elapsed_ms means the load is dominated by new terms entering the dictionary. On a large first load of an empty database, the staged loader is the faster route.

See also ​

pgRDF is released under the MIT license. Documentation built with VitePress, served via GitHub Pages.