Skip to main content
Version: V2-Next

Model-Centric Data Flow

Model Forge manages Elements as living artifacts — not as static files. Every schema has a stable global identity, is versionable, referenceable and can be projected into multiple consumable representations without modifying the original model.

The core principle: the schema is the contract — between producers (DataSources), transformations (Mappings) and consumers (DataSinks). All views are different windows onto the same contract, tailored to the respective consumer.

Model-centric data flow

Each stage adds information without destroying the previous one: the original schema in the registry remains untouched; views are ephemerally computed projections.

JSON Schema as the system language​

All entities in the system — sensor readings, road segments, observations — are modelled as JSON Schema 2020-12 documents. Every schema carries a CORE URN as its $id:

{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "urn:core:platform:civitas:element:common:GeoPoint:3ak90vqrog:1.0.0",
"title": "GeoPoint",
"type": "object",
"required": ["lat", "lon"],
"properties": {
"lat": { "type": "number" },
"lon": { "type": "number" },
"elevation": { "type": "number" }
}
}

The URN is not an opaque UUID but a readable, stable identity from which the artifact type, logical identity, version and display name are derived mechanically — see URN Format.

Composition via $ref​

Schemas reference other schemas through their CORE URN as the $ref value:

{
"$id": "urn:core:platform:civitas:element:common:TrafficObservation:eemwgr20mn:1.0.0",
"title": "TrafficObservation",
"properties": {
"location": { "$ref": "urn:core:platform:civitas:element:common:GeoPoint:3ak90vqrog:1.0.0" },
"timestamp": { "type": "string", "format": "date-time" },
"count": { "type": "integer" }
}
}

This is deliberately not a JSON-Schema-internal reference (#/$defs/…) but an external registry reference. The dependency is explicit, globally addressable and independent of the particular document: TrafficObservation can be stored in any context — the reference to GeoPoint stays stable.

The import flow​

Schema import flow

An Element import (ModelForge.importSchema) runs four steps:

  1. Validate — networknt validates the document against the JSON Schema 2020-12 meta-schema (structure). x-core-ref is a non-validating annotation here; a registry-aware step then checks that every concrete x-core-ref foreign-key target exists in the registry (ReferenceExistenceValidator, diagnostic unresolved-core-ref) — skipped when no registry is configured.
  2. Normalise — Model Forge owns identity: it mints the CORE URN (name from title, plus a disambiguator segment). The one exception is this import path — if the document already carries a real CORE-URN $id (importing an existing or externally-authored URN), that $id is kept. All $ref URNs are extracted. (createArtifact for the other artifact types never keeps a caller id — it always mints.)
  3. Store in the registry — a transactional write: upsert the artifact, insert a new artifact_version, and store the $ref URNs as artifact_reference edges (the full graph, including cycles). A content hash keeps re-imports of identical content idempotent.
  4. Update the dependency graph — forward and reverse edges in memory.

The PostgreSQL registry is the primary persistence layer, not just a cache. The dependency graph is reconstructed at backend startup from the stored artifact_reference edges across all artifacts — Elements, Mappings, Pipelines, DataSets, DataStructures, DataSources and DataSinks (ADR 054/ADR 056).

Recommendation

When importing a real, already-known CORE URN, declare it as $id in the source document — that identity is then kept. Without a $id, Model Forge mints one, which is the normal case.

Validation on two levels​

Level A — schema syntax (on import): is this JSON a valid JSON Schema? Fails on invalid type, missing required, wrong format.

Level B — reference consistency: do the artifact's references point at things that actually exist? DataSet validation checks the manifest's structure against the bundled CORE DataSet schema:

  • Dry run — validate without saving; always returns {valid, diagnostics}.
  • On save — the same check runs before persisting; errors reject the write.

x-core-ref is a typed foreign-key annotation — it marks a string field as holding the CORE URN of another artifact:

"stehtAn": {
"type": "string",
"x-core-ref": { "type": "urn:core:platform:civitas:element:common:Strasse:u8pwgr2zzg:1.0.0" }
}

It is not a validating JSON Schema keyword (it is a non-validating annotation). Instead, on Element import/update a registry-aware step checks that every concrete x-core-ref target exists in the registry (ReferenceExistenceValidator, diagnostic unresolved-core-ref); urn:core:type:<Kind> category markers and the no-registry mode are skipped. The full semantics are in the CORE-IR Reference.

The dependency graph​

In parallel to persistence, the dependency graph keeps all dependencies in memory — bidirectionally and version-precise: its nodes are versioned URNs and the queries below report concrete versioned IDs (a logical or :latest query resolves to the current version):

QueryDescription
DependenciesWhat does this schema need directly?
DependentsWho directly depends on this schema?

Transitive traversal powers the generated views internally (a bundled view embeds the whole dependency closure).

Delete protection: deleting an artifact (ModelForge.deleteArtifact) fails with ArtifactInUseException while any other artifact still references it — the exception lists the blocking dependents. This is full referential integrity: a container (a DataSet, Pipeline, DataStructure, DataSource, DataSink or Mapping) can always be deleted without touching its members, but a member cannot be deleted while a container still references it. Every edge type blocks — a grouping (datastructure-ref) protects its member just like a hard dependency does — and only an artifact's own self-references are exempt. See Composition and delete protection.

The two views​

Two qualitatively different views can be generated from the same stored schema — both on demand, without modifying the original (ADR 055).

a. Inlined view​

Goal: a single, standalone document without external dependencies.

ModelForge.getInlinedView(SchemaViewQuery)

All CORE URN $ref values are recursively replaced with the referenced schema:

{
"title": "TrafficObservation",
"properties": {
"location": {
"type": "object",
"properties": {
"lat": { "type": "number" },
"lon": { "type": "number" }
}
},
"timestamp": { "type": "string", "format": "date-time" },
"count": { "type": "integer" }
}
}

Cycle protection: if a URN reappears during inlining, the $ref node is kept.

Benefit: can be validated without a registry. Suitable for code generation, Swagger/OpenAPI, flow translators, external systems without registry access.

b. Bundled view​

Goal: bundle all transitive dependencies into a single document — as JSON Schema embedded resources under $defs. Each embedded dependency keeps its $id; CORE URN $refs stay absolute and resolve against these $ids.

ModelForge.getBundledView(SchemaViewQuery)
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "urn:core:platform:civitas:element:common:TrafficObservation:eemwgr20mn:1.0.0",
"properties": {
"location": { "$ref": "urn:core:platform:civitas:element:common:GeoPoint:3ak90vqrog:1.0.0" }
},
"$defs": {
"GeoPoint": {
"$id": "urn:core:platform:civitas:element:common:GeoPoint:3ak90vqrog:1.0.0",
"type": "object",
"properties": { "lat": { "type": "number" }, "lon": { "type": "number" } }
}
}
}

Difference from the inlined view: the types remain named, identity-bearing references and are not inlined. The document is more compact, correct for schemas with types used multiple times, and stays self-contained even for dep→dep references (chains, diamonds, cycles).

Benefit: portable, valid JSON Schema 2020-12. Can be handed to external tools, IDEs and validators without registry access. Preserves type identity.

Reading views and dependencies​

Each read is one facade call — the caller picks the projection it needs:

NeedFacade call
Raw stored documentgetArtifact(artifactId)
Bundled view (deps embedded in $defs)getBundledView(query)
Inlined view (every $ref inlined)getInlinedView(query)
Direct dependencies / dependentsdependencies(query) / dependents(query)
Mapping edges of an ElementmapsTo(query) / mappedFrom(query)

Dependencies between artifacts​

Not only schemas depend on each other — all artifact types have relationships:

The DataSet manifest is the bracketing document: it references all artifacts by their URNs but stores no content inline. A DataSet read returns the manifest as stored; each referenced artifact is fetched individually by its URN.

Self-describing artifacts: $schema in all types​

All stored artifacts carry a $schema field that declares their type — consistent with the JSON Schema convention. This makes every exported JSON file readable without context:

Artifact type$schemaIdentity
Elementhttps://json-schema.org/draft/2020-12/schema$id
DataStructurehttps://civitasconnect.digital/core-datastructure/v1id
DataSethttps://civitasconnect.digital/core-dataset/v1id
Mappinghttps://civitasconnect.digital/core/mapping/v1id
Pipelinehttps://civitasconnect.digital/core/pipeline/v1id
DataSourcehttps://civitasconnect.digital/core/datasource/v1id
DataSinkhttps://civitasconnect.digital/core/datasink/v1id

The meta-schemas themselves ship with Model Forge as bundled schema resources; a host application can expose them to its consumers if needed. All eight are described — with downloads — in the JSON Schema Reference.