Healthcare & Medical Services
21 Aug
Senior Engineer Azure Databrick
We are seeking an experienced Senior Azure Databricks Engineer to lead the design, development andoptimization of scalable and fault tolerant data solutions using pipelines and notebooks, implementingstrong data quality and governance solutions in Lakehouse architecture. You will work acrossengineering, analytics, and business teams to build robust, high-performant data workflows and ensurebest-in-class delivery on Azure and Databricks platforms.
INDICATIVE KEY RESULT AREAS (KRAs)
Platform & Infrastructure:• Oversee Databricks platform configuration, resource management, workspace structuring, andcluster optimization.• Monitor and troubleshoot performance issues across clusters, jobs, notebooks, and pipelines.• Implement governance, security, compliance and data access control using Role-Based AccessControl (RBAC) and Unity Catalog.
Pipeline Development & Architecture:• Design and implement end-to-end data pipelines using PySpark, SQL, and Delta Lake withina medallion architecture using Data factory And Data bricks.• Build real-time and batch DLT pipelines using Databricks' Delta Live Tables with a focus onreliability and scalability.• Optimize Lakehouse architecture for performance, cost-efficiency, and data integrity.• Automate data ingestion, transformation, and validation, including support for streaming(Autoloader) and scheduled workflows.• Perform data transformations, cleansing and validations using data quality rules for consistentand accurate data sets.• Manage and monitor job orchestration, ensuring efficient pipelines run and reliability.
CI/CD & DevOps:• Design and maintain CI/CD pipelines for Databricks artifacts (notebooks, jobs, libraries) usingtools such as Azure DevOps, GitHub Actions, Terraform or Jenkins.• Support trunk-based development, deployment workflows, and infrastructure-as-codepractices.• Manage version control and automated testing using Git and related DevOps practices.
Collaboration & Delivery:• Collaborate with product owners, business stakeholders, and data teams to gather requirements and translate them into technical solutions.• Drive the adoption of best practices in coding, versioning, testing, deployment, monitoring, and security.• Provide thought leadership on the best practices in Data Engineering, Architecture and Cloud Computing.
Performance Optimization:•Deliver optimized spark jobs and SQL queries for large scale data processing.•Implement partitioning, caching and indexing strategies to improve performance and scalability of big data workloads.•Conduct POCs for capacity planning and recommend appropriate infrastructure optimizations for cost effectiveness.
Documentation & Knowledge Sharing:•Created detailed documentations and review them for data workflows, SOPs, Architectural reviews etc.•Mentor junior team members and promote a culture of learning and innovation.•Promote the culture of optimization and cost saving and enable research driven development.
EDUCATION AND EXPERIENCE
Education:Bachelor's or Master’s degree in Computer Science, Engineering, Data Science, or a related field.Preferred Qualifications:•Databricks or cloud certifications (e.g., Databricks Certified Data Engineer Associate/Professional, Azure Data Engineer Associate).•Advanced expertisein PySpark and Spark DAG orchestration and optimization techniques.•Automation experience with CI/CD pipelines using Azure DevOps, Jenkins, or Octopus.•Familiarity with data mesh principles, data governance,and distributed architecture patterns.•Knowledge on observability tools, Airflow, DBT, Snowflake, Fabric, Fivetran is a plus.
Experience:Technical Expertise:•6+ years in data engineering, with a strong focus on Databricks and Azure ecosystems.•Deep hands-on experience with Data Factory, Databricks Lakehouse Architecture, Delta Lake, PySpark, and Spark job optimization.•Proficiency in Python, SQL, and optionally Scala for building scalable ETL/ELT pipelines.• Strong SQL skills are essential, with hands-on experience in SQL Server or other RDBMS platforms.• Strong experience in designing and optimizing DLT pipelines, managing assets like notebooks and libraries, and configuring Databricks Workspaces. Knowledge on observability tools, Airflow, DBT, Snowflake, Fabric, Fivetran is a plus.