Pipeline Engine Requirements
The platform needs a pipeline engine. The engine moves data from a Data source to a Data storage. On the way, it can transform the data.
A Pipeline is part of a Dataset. A Dataset is part of a Data pool. The engine runs in a Kubernetes namespace with limited rights.
This page gives the requirements. NiFi Usage Concept shows how Apache NiFi satisfies them, and which requirements are still open. ADR 047 records the decision for Apache NiFi.
Execution models
The engine must support these execution models:
- Scheduled. A cron expression starts the Pipeline.
- Event-driven. An MQTT message or a webhook call starts the Pipeline.
- Request and response. An HTTP endpoint gives a synchronous answer.
- Batch and stream. The engine must process a large set of records, and also a continuous flow of records.
Connectors
The engine must have a Connector for MQTT, for HTTP and for SQL.
It must also read and write files. Files can be text or binary. The content can be structured or unstructured.
Processing
The engine must do these operations:
- Transform data, primarily JSON.
- Apply a Mapping and validate the result.
- Write to a Data storage: SQL, HTTP, MQTT and other targets.
- Control the flow: if/then/else, switch/case and, if possible, loops.
- Give access to data through an API.
- Handle errors. It must retry, and it must move a failed record to a dead-letter path.
Performance
Scaling
The engine must scale horizontally and vertically. Two properties must scale: the number of Pipelines, and the parallelism of one Pipeline.
Expected volume
These values apply to a large city. Berlin is the reference.
| Parameter | Value |
|---|---|
| Datasets, total | up to 5,000 |
| Data pools | 200 to 300 |
| Datasets with a Pipeline | approximately 500; the other Datasets are static |
| Tables per Dataset | approximately 4 |
| Rows per table | approximately 250 |
| Users, read only | up to 20,000 |
| Users, administrative | approximately 400, which is 2 % |
| High-frequency data from IoT devices | one message every 0.5 s to 1 s |
Two values are not yet known. The first is the quantity of dynamic Datasets with sensors in a city administration. The second is the frequency and the volume of IoT use cases.
Resource efficiency
The engine must use resources efficiently. An unused Pipeline must not consume many resources.
Security
Isolation
Each Pipeline is a security domain. The Pipelines of one Dataset are in the same domain.
The engine must give three types of isolation:
- User to user. A user must not get access to the Pipeline or the data of a different user without a Permission.
- User to platform. A user must not get access to internal resources of the platform.
- Pipeline to Pipeline. A Pipeline must not get access to a different Pipeline.
Access control uses Roles, Permissions and Datasets. A Pipeline is part of a Dataset.
Confidential Datasets
A confidential Dataset needs an explicit release before a different Data pool can use it. The confidential property must stay visible everywhere.
A planned feature lets the owner of a confidential Dataset release it. The release makes a Data source that other Datasets can use.

In the example, user Lucky can use Dataset B in Data pool A. This is not sufficient. Lucky also needs a READ Permission for the yellow entities. Two Permissions are necessary: USE between the two green entities, and READ between the yellow entity and the green entity.
Threat model
| Attacker | Example | Classification |
|---|---|---|
| External attacker, not authenticated | Network access to a Pipeline endpoint | Must be prevented. Use ingress rules, network policies and authentication. |
| Authenticated user, read only | Access to the Dataset of a different user | Must be prevented. Use access control on the Pipeline and on the Dataset. |
| Authenticated user, administrative | Malicious code in an own Pipeline that reads foreign data | Must be limited. Isolate the Pipeline, or limit the available processors. |
| Platform administrator | Full access | Trusted. This is not a protection goal. |
| Advanced persistent threat | Attack on a weakness in the JVM or in the kernel | Scan the code of script nodes. Give dangerous nodes only to trusted users. |
The third attacker is the important one. An administrative user writes the Pipeline, and the engine runs it. The selected engine must show how it prevents this attack. It can isolate the process, it can remove the dangerous processors, or it can do both.
Operation
- The engine runs in a Kubernetes namespace with limited rights. It must not need a CRD or a ClusterRole.
- Helm installs the engine.
- The engine must run on premises and in an air-gapped network.
- An external registry keeps the versions of the Pipeline definitions.