Build plugins
Build plugins (also called lifecycle-hook plugins) let you run custom code at fixed points in the postprocessor’s own run — before/after each building block at each processing stage, and at a handful of run-level checkpoints (register assembly, semantic uplift, run completion, failure).
They are a different mechanism from transform plugins and
validator plugins: those run against a single building block’s
example snippets, one at a time. A build plugin instead observes or acts on the run as a whole —
stamping extra fields into register.json, triggering a notification, validating the assembled
register against an external system, syncing output somewhere.
Declaring a build plugin
plugins:
build:
- classes: [my_org.my_hooks.MyBuildHooks]
pip: git+https://github.com/example/my-bblocks-build-plugin.git
Unlike plugins.transforms/plugins.validators, which declare modules and get scanned for
matching classes by duck typing, build plugins declare classes: fully-qualified module.ClassName
strings. Each named class is imported and instantiated directly — no scanning, no discovery.
pip accepts any specifier pip install understands (package name, version constraint, Git URL,
local path); url optionally overrides the display URL recorded in register.json (derived
automatically from pip otherwise). classes can list several classes, from one or more
plugins.build entries; if more than one implements the same event, they run in declaration order,
each seeing the previous one’s output.
A build plugin is an ordinary installable Python package — no base class, decorator, or special
package metadata is needed. Its class is constructed with no arguments (or with its entry’s
config, if one is declared); it implements only the
lifecycle methods it cares about, and any event with no matching method is simply skipped.
Configuration and multiple instances
plugins:
build:
- id: strict # optional; letters, digits, _ and -
classes: [my_org.my_hooks.MyBuildHooks]
pip: git+https://github.com/example/my-bblocks-build-plugin.git
config: # optional; JSON-serializable mapping
threshold: 3
- id: lenient
classes: [my_org.my_hooks.MyBuildHooks] # same class, second instance
pip: git+https://github.com/example/my-bblocks-build-plugin.git
config:
threshold: 10
configis passed to the constructor of every class in its entry, as a single positional dict:def __init__(self, config). With noconfig(or an empty one) the class is constructed with no arguments, as before. Giving a non-emptyconfigto a class whose constructor can’t take one aborts the run. To configure classes differently, declare separate entries.configmust be JSON-serializable (string keys, noNaN; quote values like dates, which YAML would otherwise parse into non-JSON types) — this is checked up front, before any plugin is installed. Validate the contents yourself and raise from the constructor to fail the run early.configis never written toregister.json, butbblocks-config.yamlis committed to git: don’t put secrets in it. Read them from environment variables (plugins inherit the environment) or from files outside the repository.idlets the same class run as several independent instances (each(class, id)pair gets its own process; entries with the samepipstill share one virtualenv). Each pair must be unique, or the run aborts at load time. Declaring the same class twice without anidstill works (the entries share one instance, so only the first one’sconfigapplies) but logs a prominent warning and will become an error in a future release. Theidreaches the plugin ascontext['pluginId'](Nonewhen absent).config/idneed a postprocessor release that supports them (v1.1.8 or later); older images ignore the keys.
Each declared class runs in its own isolated virtualenv (created automatically under the postprocessing sandbox), the same mechanism used for transform/validator plugins.
Lifecycle events
| Method | Signature | Fires |
|---|---|---|
before_run |
before_run(self, register, context) |
Once, after the register/blocks are loaded, before any block is processed. |
before_bblock |
before_bblock(self, stage, bblock, register, context) |
Once per (stage, block), for each of the five processing stages. |
after_bblock |
after_bblock(self, stage, bblock, register, context) |
Once per (stage, block). |
after_register |
after_register(self, register, context) |
Once, after register.json’s content is fully assembled, before it’s written to disk. |
after_uplift |
after_uplift(self, register, context) |
Once, after JSON-LD/Turtle semantic uplift, before any SPARQL push. |
after_run |
after_run(self, register, context) |
Once, only on a successful run, at the very end (after a SPARQL push, if enabled). |
on_error |
on_error(self, error, register, context) |
Once, only when the run aborts. Mutually exclusive with after_run — exactly one of the two fires per run. |
stage is one of ANNOTATE, JSONLD, FINALIZE, TRANSFORMS, DOC (as a string), matching the
order the postprocessing pipeline itself loops through. Processing is stage-major, not block-major:
every block’s ANNOTATE pair fires before any block’s JSONLD pair, and so on — a plugin can’t
assume that seeing after_bblock(ANNOTATE, X, ...) means block X’s later stages are close behind.
after_bblock(FINALIZE, ...) is the one per-block event guaranteed to fire exactly once for every
block on every run (the only per-block loop that isn’t gated by --steps) — the natural anchor for
“do something once per block” plugins. At that point a block’s metadata is final except for its
documentation key, which is generator-owned and gets replaced later during the doc stage.
before_bblock/after_bblock are pure observers — nothing they return is used. after_register
is the one mutation point in the contract: if it returns a dict, that dict replaces the register
seen by the next plugin and, ultimately, what’s written to register.json. Returning anything else
leaves the register unchanged.
What each hook receives
context (present on every event) carries the run’s own settings — rootDir (the absolute
directory the other paths are relative to), pluginId (the declaring entry’s id, or null),
itemsDir, baseUrl, registerFile, steps, filter, failOnError. For per-block events at FINALIZE/DOC, it
additionally carries light (boolean): whether this block is excluded by --filter for this
stage. It isn’t present for run-level checkpoints.
bblock (per-block events only) is a plain dict, rebuilt fresh for every call: identifier, the
block’s current metadata, and urlsResolved (False through ANNOTATE/JSONLD, True from
FINALIZE onward).
register is always a plain dict, never a live object — same JSON-snapshot boundary as
transform/validator plugins. Its shape differs by event: at before_run and at every per-block
event it’s just {"bblocks": [<identifier>, ...]} (a list of identifiers — the register isn’t
fully assembled yet). At after_register, after_uplift, and after_run it’s the real, complete
register.json content.
error (on_error only) is a plain dict, not an exception object: type (exception class name),
message, traceback, and phase (one of before_run, annotate, jsonld, finalize,
transforms, doc, register, after_register, uplift). register is nullable here — a
failure during before_run, for example, has no register yet.
Failure semantics
- Run-level checkpoints (
before_run,after_register,after_uplift,after_run): any exception always aborts the run, regardless of--fail-on-error. - Per-block events (
before_bblock/after_bblock): follow the same rule as the rest of the pipeline — with--fail-on-errorset, an exception aborts the run; otherwise it’s logged and processing continues, with the block left intact. A non-fatal hook failure leaves only a log line — no marker is written intoregister.jsonor the block’s metadata. on_errorfires exactly once, only when the whole run aborts, and never together withafter_run. It doesn’t fire for non-fatal per-block errors, for a failed SPARQL push (handled internally and still counted as a successful run), or for failures before plugins are even loaded (e.g. a brokenbblocks-config.yaml).
If you rely on a hook actually running, run with --fail-on-error — otherwise a failing hook is
easy to miss.
Minimal example
class MyBuildHooks:
def before_run(self, register, context):
print(f"starting run over {len(register.get('bblocks', []))} block(s)")
def after_register(self, register, context):
# the one mutation point: stamp a custom field into the register
result = dict(register)
result['x-myBuildPlugin'] = {'bblockCount': len(result.get('bblocks', []))}
return result
def on_error(self, error, register, context):
print(f"run failed in phase {error['phase']}: {error['message']}")
A fuller, real-world example implementing every event is
bblocks-build-plugin-sample
(SampleBuildHooks), which stamps a processing timestamp onto the register and every block via
after_register.
A build plugin that stamps a new field/document onto a bblock pairs naturally with a
tab plugin on the viewer side: the build plugin emits the data into
json-full, the tab plugin renders it as its own tab. The two are independent mechanisms, though —
a tab plugin can just as validly key off existing bblock/register metadata with no build plugin
involved at all.
Plugin metadata in the register
Approved build plugins are recorded in register.json under buildPlugins (classes, pip
specifier(s), URL) — parallel to the existing transformPlugins/validatorPlugins keys. This
matters more here than for transform/validator plugins, since after_register can silently rewrite
the register: a consumer can see from buildPlugins that plugin code had the opportunity to edit it.
See Security for the shared permission/sandboxing model across plugin kinds.