Working as a Data Engineer Consultant, you will:
-
Design and implement data ingestion pipelines connecting external APIs and business systems with the Databricks data platform.
-
Develop reliable processes for retrieving and processing historical data, including mechanisms for rate limiting, checkpointing, retries and safe reprocessing.
-
Build and maintain the Bronze layer of the data platform, covering data standardisation, typing, historical tracking and document storage.
-
Develop processes for linking digital documents with corresponding business records while ensuring data consistency and idempotent processing.
-
Use Databricks capabilities to classify documents and extract structured information from PDFs and images, including confidence scoring and workflows for manual review.
-
Contribute to the development of the Silver layer using appropriate entity modelling, CDC and incremental processing patterns.
-
Implement deterministic entity resolution and source-priority logic across multiple data sources.
-
Apply data quality controls within the existing framework and ensure that quality results are properly stored, monitored and available for governance purposes.
-
Support the development of analytical and machine-learning data products, including feature and metric stores and batch scoring workflows.
-
Work with MLflow and related processes for model registration and promotion where required.
-
Improve automation, monitoring and reliability across data pipelines and contribute to CI/CD processes.
-
Collaborate with other technical stakeholders to deliver agreed project milestones within a defined consulting scope.