ADR 061: portal-backend integration converts XSD at first import
Date: 2026-08-27
Status: Accepted
Decision Makers: @derlinne, @luckey
Context
Model Forge supports XSD and JSON Schema as equal formats per version (ADR 058): an XSD-authored version keeps its original and additionally exposes a derived JSON Schema, generated on read and cacheable.
For the portal-backend integration that either/or is a liability. Carrying two possible
content types through host orchestration, config-adapter and the frontend would
spread format handling across layers that must not interpret model payloads at all
(see portal-backend Integration),
and it would move the conversion — and any conversion diagnostics — onto an arbitrary
later read instead of the import, where a human is present.
Decision
On the first import of an XSD, the portal-backend integration triggers the JSON Schema
conversion immediately and persists it as that version's generated representation.
All model processing above the Model Forge boundary works on JSON Schema only.
The submitted XSD remains stored as the version's authored original and stays
retrievable — including through the portal-backend API. Getting back to the original input
is a supported operation, not an internal detail.
Rationale
- One content type for everything that processes models, so no format branch appears in a layer that is not allowed to understand payloads anyway.
- Conversion cost and conversion diagnostics land at import time.
- Model Forge needs no new capability: this is the eager use of the
generatedrepresentation it already maintains. - The original is what makes a conversion auditable and a corrected conversion reproducible, so it must stay reachable from the outside — not only from inside the registry.
Mechanism
- The import stores the XSD as the authored representation (
generation = 'stored',primary_format = xsd) and, in the same transaction, the converted JSON Schema asgeneration = 'generated'. - The version decision is made on content, never on a converter version. An import creates a new version when the submitted XSD or the converted JSON Schema differs from the current version's; identical input and identical result is a no-op. Nothing is ever rewritten in place, so no existing version changes underneath its consumers.
- A converter bugfix therefore lands by re-importing the unchanged XSD: the result differs, so the import is a new version. A converter change that alters nothing about this document produces no version. No converter version is recorded, compared or reasoned about — the content answers the question.
- The comparison is structural, on the parsed documents — not on serialised bytes.
Object key order carries no meaning in JSON Schema, so a mere reserialisation must not
register as a change. Array order does count (
oneOf,enum,required), which is correct: a reorderedoneOfis a real change. content_hashkeeps its existing job — a byte hash of what a read serves, for ETags and lost-update detection. It is deliberately not the version predicate: it is byte-stable but not canonical, so two equal documents with different key order hash differently.- This requires the converted schema to be persisted: today it is derived on every read and never stored, so there is nothing to compare against.
- Reads of an XSD-imported version return JSON Schema by default; the
xsdrepresentation is served verbatim on explicit request, through Model Forge and through theportal-backendAPI.
Consequences
- Model Forge needs two changes, both inside the existing storage model. Today the
idempotency check compares only the authored content of the current version in the
same format, and
generation = 'generated'is never written — the converted schema is derived on every read. So a bugfixed converter currently cannot land at all: the XSD is byte-identical, the write is a no-op, and the improved conversion is never stored. The import must therefore (1) persist the converted schema as the version'sgeneratedrepresentation and (2) extend the version decision to that representation's hash. No new column, no new format, no change to the facade's shape. - Because the decision is content-based, re-importing an unchanged XSD through an unchanged converter stays a no-op — no version spam from repeated imports.
- The converter must introduce no per-run variation into the compared document. This
already holds and is deliberate:
XsdToJsonSchemaConvertersorts the type iteration becausegetSchemaTypes()is HashMap-backed and "not stable across JVM runs",required/oneOffollow XSD document order, and disambiguators are derived from the name rather than minted randomly. Structural comparison keeps this robust — it survives a Jackson upgrade or a formatting change that a byte hash would not. - Because the conversion is persisted rather than derived on demand, it has a stored content hash — so a JSON Schema read of an XSD-imported Element is conditional-request capable, which a non-persisted derived representation is not.
- Serving the original XSD through the
portal-backendAPI is envelope-level passthrough: the host forwards bytes it does not parse, which is on the correct side of the content cut. It must not become a second place where format decisions are made. - Conversion coverage stays the quality gate for XÖV standards. The known converter
limitation for classic (non-CORE-URN) XSD namespaces — a placeholder
$refinstead of the imported target — remains a defect in a derived representation while the original stays authored. Repairing it in the converter and re-importing the unchanged XSD yields a new version, because the result changed: exactly the behaviour the content-based decision is there for. - An Element imported as XSD stays renamable and versionable exactly as before; the authored format of a version is unchanged by this decision.
- The
xsdread format is neither removed nor deprecated.
See also
- ADR 058: Format-agnostic Elements with versioned representations
- ADR 048: JSON Schema 2020-12 as the canonical modelling format
- portal-backend Integration & Division of Labor
- XML/XSD Integration