Verbose ingest statistics
A JSONB report from an ingest run: which loader path ran, how many triples went in, how many lines were skipped, how much dictionary work it took, and where the time went.
What it does
pgrdf.parse_turtle_verbose(content TEXT, graph_id BIGINT, base_iri TEXT DEFAULT NULL) → JSONB
pgrdf.load_turtle_verbose(path TEXT, graph_id BIGINT, base_iri TEXT DEFAULT NULL,
bulk_load BOOLEAN DEFAULT FALSE) → JSONBThese load exactly like parse_turtle and load_turtle, but return a report instead of a bare count. parse_trig, parse_nquads and load_turtle_streaming return the same report; the quad formats add a "graphs" array.
Why you'd use it
- Project managers — measure ingest cost on your own data and plan capacity from observed numbers.
- Data scientists — check that a fast path actually ran before committing to a large load.
- Operators — see whether time goes on parsing, dictionary lookups or inserting.
Example
sql
SELECT pgrdf.add_graph('http://example.org/people');
SELECT jsonb_pretty(pgrdf.parse_turtle_verbose('
@prefix ex: <http://example.org/> .
@prefix foaf: <http://xmlns.com/foaf/0.1/> .
ex:alice a foaf:Person ; foaf:name "Alice" ; foaf:knows ex:bob .
ex:bob a foaf:Person ; foaf:name "Bob" .
', pgrdf.graph_id('http://example.org/people')));json
{
"path": "combined",
"triples": 5,
"parse_skipped": 0,
"quad_batches": 1,
"dict_db_calls": 9,
"dict_cache_hits": 8,
"shmem_cache_hits": 0,
"dict_terms": 0,
"windows": 0,
"defer_index": false,
"parse_ms": 0.04,
"dict_ms": 0.70,
"insert_ms": 0.19,
"resolve_ms": 0.0,
"index_ms": 0.0,
"elapsed_ms": 0.94
}(Keys reordered for reading, timings rounded; timings vary by machine.)
Reading the report
| Key | Meaning |
|---|---|
path | Which loader ran: combined (the standard parser), parallel_bulk (bulk_load => true) or streaming (load_turtle_streaming). |
triples | Triples loaded. |
parse_skipped | Lines the N-Triples fast paths skipped as unreadable. Always check this after a bulk_load or streaming run. |
quad_batches | Number of batched inserts used to write the triples. |
dict_db_calls | Term lookups that went to the dictionary table (mostly new terms). |
dict_cache_hits | Term lookups answered from memory without a table lookup. |
shmem_cache_hits | Term lookups answered by the shared-memory term cache. |
dict_terms, windows, defer_index | Work done by the bulk and streaming paths: terms added in memory, windows processed, whether indexes were rebuilt after the load. |
parse_ms, dict_ms, insert_ms, resolve_ms, index_ms | Time per phase, in milliseconds. |
elapsed_ms | Total time inside the call, in milliseconds. |
Two readings worth knowing:
triples: 0withparse_skipped > 0afterbulk_load => truemeans the input wasn't N-Triples, so nothing loaded. Load Turtle withoutbulk_load.dict_msclose toelapsed_msmeans the load is dominated by new terms entering the dictionary. On a large first load of an empty database, the staged loader is the faster route.
See also
- Bulk ingest — the loader paths the report describes.
- Shared-memory term cache.
pgrdf.stats()— server-wide counters.