Job Description
Role: Senior Data Engineer, Azure Databricks
Location: Onshore (Remote)
Please note: We need pre-vetted profiles from your end as this is Senior level position for Data Engineer.
Please share the resumes with LinkedIn profiles only and preferably from the pharmacy domain and share the profiles directly with me.
Please share 1 strong profile to crack the interview.
Purpose
The Senior Data Engineer builds, operates, and owns the enterprise data warehouse and core data pipelines on Customer's Databricks platform. This role is responsible for EDW architecture, ingress patterns into Bronze, transformation logic from Bronze to Silver, and Gold harmonization rules. The Senior Data Engineer significantly contributes to the platform's evolution to a simplified medallion architecture with event-driven, low-latency ingestion. This role serves as a senior practitioner within the Data Engineering track, setting the technical bar for pipeline quality, documentation discipline, and reliability, as well as mentoring Data Engineers as the team scales.
Responsibilities
EDW Engineering & Ownership
- Designs, builds, and operates enterprise data warehouse pipelines and data models across the Databricks medallion architecture (Bronze, Silver, Gold), including Silver conformance logic and Gold harmonization and survivorship rules.
- Assumes named ownership of assigned EDW domains as production systems: data model integrity, pipeline reliability, ingestion SLAs, and incident response.
- Leads structured documentation of knowledge for assigned domains, absorbing architecture, transformation logic, and feed SQL.
- Executes data model simplification work, consolidating legacy transformation layers into the target architecture with validated output parity.
Pipeline Development & Streaming
- Develops and maintains production Delta Lake pipelines, evolving ingestion from batch watermark patterns to event-driven architectures using Zerobus Ingestion and Spark Declarative Pipelines for high-SLA source systems.
- Implements row-level reconciliation, validation checkpoints, and quality gates at ingestion, transformation, and delivery layers, operationalizing governance-defined data quality dimensions in partnership with QA.
- Builds Gold-layer tables and promotion logic that support Unity Catalog metric views, lineage capture, and access controls in partnership with Analytics Engineering and Data Governance.
- Optimizes pipeline performance, latency, and compute cost across all layers of the platform.
Documentation, Quality & Observability
- Documents architecture decisions, data models, pipeline logic, and operational runbooks to team standards, ensuring no critical platform capability depends on a single point of failure.
- Operates pipeline observability tooling: monitors anomaly detection, triages data reliability incidents, and drives root cause remediation for assigned domains.
- Implements data contract validation (ODCS or equivalent) in pipelines supporting external partner feed domains, ensuring contract failures halt delivery before reaching consumers.
Mentorship & Cross-Functional Collaboration
- Mentors Data Engineers on Databricks development patterns, SQL and PySpark craft, and documentation discipline; reviews pull requests in Azure DevOps & Git to maintain code quality.
- Partners with Data Governance on lineage capture and metadata standards, with QA on Tier 1 and Tier 2 pipeline test suite development, and with Partner Data Services on feed engineering requiring pipeline work.
- Works from structured requirements and acceptance criteria entering through the Informatics intake process, and flags requirements gaps before build work begins.
Required Qualifications
- 6+ years of progressive data engineering experience, including production ownership of an enterprise data warehouse or large-scale transformation pipelines with defined SLAs.
- Deep, hands-on expertise with Databricks in production environments: Delta Lake, medallion architecture, Unity Catalog, and pipeline performance optimization.
- Demonstrated data warehousing depth: dimensional and harmonized data modeling, conformance and survivorship logic, and operating a warehouse as a production system.
- Advanced SQL and strong PySpark and Python proficiency, sufficient to build, review, and optimize complex transformation logic independently.
- Proven experience with Azure data services: Azure Databricks, Azure Data Lake Storage, Azure Data Factory, and CI/CD practices in Azure DevOps.
- Demonstrated ability to absorb complex, under-documented systems through structured knowledge transfer and produce documentation that makes that knowledge durable and transferable.
- Strong communication and collaboration skills; able to work directly with QA, Governance, Informatics, and business-facing teams without an intermediary.
Preferred Qualifications
- Healthcare, specialty pharmacy, or regulated industry experience, with exposure to clinical or operational data environments.
- Production experience with event-driven ingestion: Azure Event Hubs, Kafka, or equivalent, consumed via Structured Streaming.
- Familiarity with data contract standards (ODCS or equivalent) and pipeline observability platforms (Monte Carlo or similar).
- Familiarity with enterprise data catalog tooling (Atlan, Collibra, or equivalent) and designing pipelines as catalogued, governed assets.
- Experience leading through influence: mentoring engineers, setting standards, and driving adoption without formal management authority.
- Bachelor's degree in Computer Science, Data Engineering, Information Systems, or a related field, or equivalent experience
Job Tags
Full time, Contract work, Remote work