Job description
We are seeking an experienced Data Engineer to design, build, and maintain our cloud-based data infrastructure on Azure and Databricks. This role is pivotal in ensuring our data platform is scalable, secure, and optimised for advanced analytics and reporting.
Key Responsibilities
Platform Development & Maintenance
• Design and implement robust data pipelines using Azure Data Factory, Databricks, and related Azure services.
• Build ETL/ELT processes to transform raw data into structured, analytics-ready formats.
• Optimise pipeline performance, ensure data quality, and maintain high availability of data services.
Infrastructure & Architecture
• Architect and deploy scalable data lake solutions using Azure Data Lake Storage.
• Implement data governance and security measures aligned with compliance standards.
• Use Infrastructure-as-Code tools (e.g., Terraform) for reproducible deployments.
• Define best practices for data organisation, partitioning, and storage optimisation.
Databricks Development
• Develop and optimise jobs within Databricks, leveraging Delta Lake for reliable storage and ACID transactions.
• Implement medallion architecture (bronze, silver, gold layers) and reusable notebooks.
• Optimise cluster configurations for cost and performance.
• Establish CI/CD pipelines for Databricks deployments.
Monitoring & Operations
• Implement monitoring and alerting solutions using Azure Monitor, Log Analytics, and Databricks tools.
• Troubleshoot performance issues, optimise resource utilisation, and ensure SLAs are met.
• Define disaster recovery procedures and backup strategies.
Collaboration & Documentation
• Partner with data scientists, analysts, and business stakeholders to deliver tailored data solutions.
• Document technical architectures, data flows, and operational procedures for knowledge sharing.
Essential Skills & Experience
• 5+ years of experience with Azure Cloud services (Data Factory, Data Lake Storage, SQL Database, Synapse Analytics).
• Strong hands-on experience with Databricks, Delta Lake, and cluster management.
• Proficiency in SQL and Python for data manipulation and pipeline development.
• Experience with Git/GitHub and CI/CD practices.
• Solid understanding of data modelling concepts and data lake architectures.
• Knowledge of data governance, security, and compliance requirements.
Desirable Skills
• Experience with Terraform or other Infrastructure-as-Code tools.
• Familiarity with Azure DevOps or similar CI/CD platforms.
• Exposure to data quality frameworks and testing approaches.
• Azure certification (e.g., Azure Data Engineer Associate).
• Databricks certification.