Skip to main content
Version: V2-Next

ADR 061: portal-backend integration converts XSD at first import

Date: 2026-08-27

Status: Accepted

Decision Makers: @derlinne, @luckey

Context​

Model Forge supports XSD and JSON Schema as equal formats per version (ADR 058): an XSD-authored version keeps its original and additionally exposes a derived JSON Schema, generated on read and cacheable.

For the portal-backend integration that either/or is a liability. Carrying two possible content types through host orchestration, config-adapter and the frontend would spread format handling across layers that must not interpret model payloads at all (see portal-backend Integration), and it would move the conversion — and any conversion diagnostics — onto an arbitrary later read instead of the import, where a human is present.

Decision​

On the first import of an XSD, the portal-backend integration triggers the JSON Schema conversion immediately and persists it as that version's generated representation. All model processing above the Model Forge boundary works on JSON Schema only.

The submitted XSD remains stored as the version's authored original and stays retrievable — including through the portal-backend API. Getting back to the original input is a supported operation, not an internal detail.

Rationale​

  • One content type for everything that processes models, so no format branch appears in a layer that is not allowed to understand payloads anyway.
  • Conversion cost and conversion diagnostics land at import time.
  • Model Forge needs no new capability: this is the eager use of the generated representation it already maintains.
  • The original is what makes a conversion auditable and a corrected conversion reproducible, so it must stay reachable from the outside — not only from inside the registry.

Mechanism​

  • The import stores the XSD as the authored representation (generation = 'stored', primary_format = xsd) and, in the same transaction, the converted JSON Schema as generation = 'generated'.
  • The version decision is made on content, never on a converter version. An import creates a new version when the submitted XSD or the converted JSON Schema differs from the current version's; identical input and identical result is a no-op. Nothing is ever rewritten in place, so no existing version changes underneath its consumers.
  • A converter bugfix therefore lands by re-importing the unchanged XSD: the result differs, so the import is a new version. A converter change that alters nothing about this document produces no version. No converter version is recorded, compared or reasoned about — the content answers the question.
  • The comparison is structural, on the parsed documents — not on serialised bytes. Object key order carries no meaning in JSON Schema, so a mere reserialisation must not register as a change. Array order does count (oneOf, enum, required), which is correct: a reordered oneOf is a real change.
  • content_hash keeps its existing job — a byte hash of what a read serves, for ETags and lost-update detection. It is deliberately not the version predicate: it is byte-stable but not canonical, so two equal documents with different key order hash differently.
  • This requires the converted schema to be persisted: today it is derived on every read and never stored, so there is nothing to compare against.
  • Reads of an XSD-imported version return JSON Schema by default; the xsd representation is served verbatim on explicit request, through Model Forge and through the portal-backend API.

Consequences​

  • Model Forge needs two changes, both inside the existing storage model. Today the idempotency check compares only the authored content of the current version in the same format, and generation = 'generated' is never written — the converted schema is derived on every read. So a bugfixed converter currently cannot land at all: the XSD is byte-identical, the write is a no-op, and the improved conversion is never stored. The import must therefore (1) persist the converted schema as the version's generated representation and (2) extend the version decision to that representation's hash. No new column, no new format, no change to the facade's shape.
  • Because the decision is content-based, re-importing an unchanged XSD through an unchanged converter stays a no-op — no version spam from repeated imports.
  • The converter must introduce no per-run variation into the compared document. This already holds and is deliberate: XsdToJsonSchemaConverter sorts the type iteration because getSchemaTypes() is HashMap-backed and "not stable across JVM runs", required/oneOf follow XSD document order, and disambiguators are derived from the name rather than minted randomly. Structural comparison keeps this robust — it survives a Jackson upgrade or a formatting change that a byte hash would not.
  • Because the conversion is persisted rather than derived on demand, it has a stored content hash — so a JSON Schema read of an XSD-imported Element is conditional-request capable, which a non-persisted derived representation is not.
  • Serving the original XSD through the portal-backend API is envelope-level passthrough: the host forwards bytes it does not parse, which is on the correct side of the content cut. It must not become a second place where format decisions are made.
  • Conversion coverage stays the quality gate for XÖV standards. The known converter limitation for classic (non-CORE-URN) XSD namespaces — a placeholder $ref instead of the imported target — remains a defect in a derived representation while the original stays authored. Repairing it in the converter and re-importing the unchanged XSD yields a new version, because the result changed: exactly the behaviour the content-based decision is there for.
  • An Element imported as XSD stays renamable and versionable exactly as before; the authored format of a version is unchanged by this decision.
  • The xsd read format is neither removed nor deprecated.

See also​

  • ADR 058: Format-agnostic Elements with versioned representations
  • ADR 048: JSON Schema 2020-12 as the canonical modelling format
  • portal-backend Integration & Division of Labor
  • XML/XSD Integration