Information Technology & Software
29 Sept
Azure Databricks Engineer
Azure Databricks Engineer
Job Role & Responsibilities
Platform & Infrastructure:
Support Databricks workspace configuration, compute resources, libraries, jobs, and cluster settings in accordance with established platform standards.
Monitor and troubleshoot performance and reliability issues across clusters, jobs, notebooks, and pipelines; document findings and escalate complex issues as needed.
Implement approved governance, security, compliance, and data-access controls using Role-Based Access Control (RBAC) and Unity Catalog.
Pipeline Development & Architecture:
Develop and maintain end-to-end data pipelines using Azure Data Factory, PySpark, SQL, and Delta Lake within a medallion architecture.
Build and support batch and real-time Delta Live Tables (DLT) pipelines with a focus on reliability, scalability, and recoverability.
Apply practical Lakehouse optimization techniques to improve performance, cost efficiency, and data integrity.
Automate data ingestion, transformation, and validation, including streaming with Auto Loader and scheduled workflows.
Perform data transformations, cleansing, reconciliation, and validation using defined data-quality rules to produce consistent and accurate datasets.
Monitor job orchestration, investigate failures, support incident resolution, and implement fixes that improve pipeline reliability.
CI/CD & DevOps:
Develop and maintain CI/CD workflows for Databricks artifacts, including notebooks, jobs, and libraries, using Azure DevOps, GitHub Actions, Terraform, Jenkins, or similar tools.
Follow established trunk-based development, deployment, automated-testing, and infrastructure-as-code practices.
Use Git and related DevOps practices for version control, peer review, testing, and controlled promotion across environments.
Collaboration & Delivery:
Collaborate with product owners, business stakeholders, and data teams to understand requirements and translate them into practical technical solutions.
Apply established practices for coding, versioning, testing, deployment, monitoring, documentation, and security.
Participate in code reviews, testing, release activities, production support, and root-cause analysis with guidance from senior engineers when needed.
Performance Optimization:
Tune Spark jobs and SQL queries for efficient processing of large datasets.
Apply appropriate partitioning, caching, file-layout, and indexing techniques to improve workload performance and scalability.
Review job metrics and resource usage, identify optimization opportunities, and implement approved performance and cost improvements.
Documentation & Knowledge Sharing:
Create and maintain clear technical documentation for data workflows, support procedures, runbooks, configurations, and standard operating procedures (SOPs).
Share troubleshooting findings, reusable code, and implementation knowledge with team members.
Contribute to continuous improvement by identifying recurring issues, automation opportunities, and practical cost-saving enhancements.
Specific Expertise Required
2-5 years of data-engineering experience, including hands-on work with Azure and Databricks technologies.
Hands-on experience with Azure Data Factory, Databricks Lakehouse architecture, Delta Lake, PySpark, and Spark job performance tuning.
Proficiency in Python and SQL for building and supporting scalable ETL/ELT pipelines; Scala experience is optional.
Strong SQL skills, with hands-on experience using SQL Server or another relational database management system (RDBMS).
Experience developing and supporting DLT pipelines and managing Databricks assets such as notebooks, jobs, libraries, and workspace resources.
Working knowledge of Unity Catalog, RBAC, data-quality controls, pipeline monitoring, troubleshooting, and production support practices.
Solid foundation in data-warehousing principles and dimensional data modeling for reporting and analytics.
Experience using Git and participating in CI/CD, automated testing, and controlled deployment processes.
Preferred Qualifications
Databricks or Azure certifications, such as Databricks Certified Data Engineer Associate or Professional, or Microsoft Azure data-engineering certifications.
Additional experience with PySpark, Spark DAG analysis, Structured Streaming, Auto Loader, and workload optimization techniques.
Automation experience with CI/CD pipelines using Azure DevOps, GitHub Actions, Jenkins, Octopus, Terraform, or similar tools.
Familiarity with data mesh principles, data governance, and distributed architecture patterns.
Exposure to observability tools, Airflow, dbt, Snowflake, Microsoft Fabric, or Fivetran is a plus.
Educational Qualifications
Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or a related field.