Data Architect
Role overview
Your mission: You define how data is structured, processed, and managed across the Platform. Your goal is to establish standards, integrate data, and ensure that data can be consistently used and reused.
Why your role is vital: You are the "Architect". You create the foundation that enables all data-related work. Without your structures, standards, and integrations, DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. cannot be created, processed, or reused.
Defining your scope: You work across the Platform on Data poolsData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices)., Data structures, Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network., Pipelines, and APIs. You define standards and enable others to create, manage, and use data-related elements.
Your core responsibilities
- Define standards and guidelines: You define Data structures, metadata standards, and reusable patterns across the tenantTenantAn isolated organizational partition that owns Data pools, Datasets, Users, Groups, and Roles. All access rules exist within their Tenant, and the Tenant is the widest Scope of a Role. Currently, one Tenant corresponds to the Platform..
- Manage data architecture: You create and maintain Data structures, Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network., and pipelines across the Platform.
- Organize data domains: Structure related DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. using Data poolsData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices). and define shared governance where appropriate
- Enable data processing: You ensure data can be ingested, transformed, stored, and provided consistently.
- Support the organization: You work across domains and support other Roles in building and using data-related elements.
Outside your scope
It is not your responsibility to release DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. or govern access policies.
If you are responsible for reviewing and releasing DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions., the Data Owner or Data Gatekeeper role is likely the right role for you.
To work effectively with data-related elements, you need the correct permissions.
The access logic: Access is always the result of a User being assigned to a Group, and that Group being assigned a Role at a specific Scope.
Scopes: Roles apply either at Platform level, on a Data poolData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices)., or on a specific data-related element such as a DatasetDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions., Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network., or Data structure.
To create and manage data-related elements, your Group typically needs a Data Role with create Permissions at Platform scope.
→ Deep Dive Authorisation
Typical tasks
Your work in CIVITAS/CORE focuses on designing and managing data-related elements across the Platform:
- Design Data structures: Create and version Data structures based on defined standards
- Register Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network.: Connect systems and ensure data is ingested correctly
- Organize DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. using Data poolsData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices).: Provide an additional scope for access management
- Build pipelines and configure APIs: Define how data is ingested, transformed, and prepared for use
- Prepare DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. for further use: Ensure DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. are complete and ready for review by Data Owner or Data Gatekeeper
Your first steps
To start working as a Data Architect in CIVITAS/CORE:
- Check that your Group has a Data Role with create Permissions at Platform scope
- Explore existing Data structures and Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. in the Platform
- Create your first Data structure
- Create your first Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. and link it to the Data structure
- Create your first DatasetDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. and assign it to a Data poolData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices). (recommended)
- Build your first Pipeline
- Configure APIs and manage access permissions
Whenever possible, manage access at Data poolData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices). level. Use DatasetDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions.-specific permissions only when individual DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. require different access rules.
Best practices & avoiding mistakes
- Work with reusable Data structures: Design Data structures so they can be reused across multiple DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions.
- Keep pipelines clear and maintainable: Avoid overly complex pipeline logic. Keep transformations understandable
- Separate responsibilities: Focus on building and preparing data. Data Stewards handle the detailed configuration of DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions., while Data Owners and Gatekeepers are responsible for release decisions
- Use Data poolsData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices). to organize related DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions.: This simplifies access management and Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. availability
- Ensure proper access setup: Make sure your Group has the required platform-wide Data Role to create data-related elements
- Enable Data Stewards through access: Assign Data Roles to Groups on the relevant DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions., Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network., and Data structures so Data Stewards can work on them.
When working with pipelines, make sure access is assigned across all involved data-related elements.
If a DatasetDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. uses a Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. or Data structure, the required Group + Role assignment must exist on each of them. When using Data poolsData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices)., remember that permission inheritance does not automatically grant access to referenced Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. or Data structures.
Otherwise, your collaborators may not be able to select Data sourcesData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. or Data structures and configure or run pipelines.
Your permissions
| Permissions / Data-related elements | READ | CREATE | UPDATE | DELETE | RELEASE |
|---|---|---|---|---|---|
| DatasetDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. | 🟢 | 🟢 | 🟢 | 🟢 | |
| ↳ Payload | |||||
| Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. | 🟢 | 🟢 | 🟢 | 🟢 | |
| Data structure | 🟢 | 🟢 | 🟢 | 🟢 | |
| Data poolData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices). | 🟢 | 🟢 | 🟢 | 🟢 |
Key terms to know
To work effectively as a Data Architect, review these terms in our Glossary:
- Data structure: Defines the schema and structure of data
- Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network.: Represents the origin of data.
- DatasetDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions.: A structured collection of data prepared for use
- Pipeline: Defines how data is loaded, transformed, and stored
- Data Role: A Role that grants access to data-related elements
- Scope: Defines where a Role applies, either on Platform level or on a specific data-related element
- Data poolData poolA governed, centralized collection of Datasets that are managed together within the Platform. It provides a shared place to organize, discover, and access data, including associated metadata, ownership, and access permissions. With Data pools users can cluster their Datasets according to their organizational structure (e.g., by Departments, Offices).: Organizes related DatasetsDatasetA data-related element that contains processed data and makes it available for consumption. A Dataset is populated via Pipelines and carries Metadata and access permissions. and provides an additional scope for access management and Data sourceData sourceA data-related element that represents the origin of data. It defines how data is connected, accessed, and ingested into the Platform, such as an external database or sensor network. availability