Community Edition / Docs / Architecture

One broker, a handful of services, a directory

The whole stack is a compose file. Here is what runs, which port it listens on, and why none of it needs a cloud account.

Every arrow is an HTTP request on localhost. No queue, no bucket, no serverless function.
ServicePortRole
web8080Walkthrough UI, JSON-LD context host, same-origin proxy to every service.
scorpio9090NEC Scorpio NGSI-LD broker, upstream image, in-memory profile, PostGIS behind it.
usdm8101Turns a USDM-style study definition into broker entities.
projection8102Reads broker state, writes Dataset-JSON trial-design domains to the lake.
subscriptions8103Receives broker notifications and triggers a transform run.
transform8104sdtm.oak in R. Regenerates SDTM from current broker state.
lake8105Three HTTP endpoints over a directory; DuckDB for SQL.
adaptive8106Natural-language statement to entity proposal. Model call or offline stub.
bridge8107Optional rclone connectors, in and out. Off by default.

The broker never needed the cloud

NEC Scorpio ships its messaging layer as first-class build variants, visible in its release tags: -java (in-memory), -kafka, -sqs, -amqp, -mqtt. The in-memory build runs every broker service in one JVM with in-process channels. The AWS coupling in Garnet-style deployments came from choosing the -sqs variant to span Fargate containers; it was packaging, not a broker requirement. This stack runs the upstream image, unmodified.

Three replacements

Subscriptions instead of SQS/SNS. One NGSI-LD subscription covers the execution entity types. The broker POSTs notifications straight to a listener; bursts coalesce into a single idempotent transform run that regenerates SDTM from current broker state.

A directory instead of S3, DuckDB instead of Athena. Everything the pipeline produces is a file you can open: Dataset-JSON design domains, raw EDC-style extracts, SDTM datasets. DuckDB gives SQL over those files with zero infrastructure.

Compose instead of Fargate. Single node is the point. Five of the seven services are standard-library Python and build with no network access at all.

Swapping the lake

The lake sits behind a deliberately small contract: writers drop files, readers use three HTTP endpoints. That makes it the natural place to diverge for production: keep the files and add Postgres as a query replica, sync to S3 and query with Athena, or run Trino, ClickHouse, or MinIO. Recipes with the exact steps are in docs/lake-alternatives.md.

The vocabulary is a plugin

Entities are typed with the cr-domain of The Ontology Project: an open, provenance-native reference model for the clinical-research lifecycle, the first anchored T.O.P. domain. Nothing in the engine is clinical-specific: the broker is standard NGSI-LD, the context is a served JSON-LD file, and the projections are readers of broker state. Anchoring another domain (healthcare delivery, commercialization, claims, manufacturing, supply chain) means a new vocabulary and new projections over the same architecture. The NGSI-LD context is vendored and served locally, so the whole stack runs with no internet access.

Adaptive layer in this stack

The adaptive service turns a natural-language clinical statement into an NGSI-LD entity proposal. POST /construct proposes entities; POST /commit executes only when the PNE_ADAPTIVE_EXECUTE gate is enabled, and every commit attempt keeps a local decision record. With an API key it calls the model once; without one it uses a deterministic offline stub. Broker-native temporal queries cover provenance and time for entities this stack writes.