EFutures logo
1
applicant

Senior Data Engineer - Azure Databricks / PySpark

EFutures    Colombo • Full-time

Job Description

Join our team and work on exciting data and MDM initiatives building scalable data pipelines and data solutions for enterprise systems.

Key Responsibilities

  • Develop ingestion pipelines for REST APIs, database/CDC and file-based sources.
  • Build PySpark/Spark transformations and Delta Lake processing on Azure Databricks.
  • Implement Bronze/raw, standardised/conformed and governed data patterns as defined by architecture.
  • Develop reusable standardisation, validation and DQ processing before MDM.
  • Implement incremental loading, schema validation/evolution and restart/recovery patterns.
  • Create source-to-canonical transformations while retaining lineage/provenance.
  • Optimise Spark jobs for performance, cost and operational reliability.
  • Write automated unit/data-quality tests and participate in peer review.
  • Integrate Databricks outputs/inputs with Profisee and downstream Azure services.
  • Support production monitoring, troubleshooting and runbook creation.

Must-Have Experience & Skills

  • Overall 7+ years of experience and 5+ years data engineering with strong production PySpark/Spark experience.
  • Strong Azure Databricks, Delta Lake, SQL and ADLS Gen2.
  • REST/API ingestion and database integration experience.
  • Incremental/CDC processing and schema-evolution experience.
  • Strong data-quality/testing practices and Git/CI/CD.
  • Experience handling large, messy enterprise datasets.

Preferred/Nice-to-Have

  • Databricks certification.
  • MDM/entity-resolution project exposure.
  • Azure Data Factory or equivalent orchestration.
  • Exposure to ML-based entity resolution / fuzzy matching (PySpark MLlib).
  • Purview/Unity Catalog governance exposure.

Employment Details

  • Employment Type: Part Time (2 Years Contract)
  • Work Arrangement: Hybrid (3 to 4 days onsite)
  • Working Days: Monday to Friday
  • Working Hours: 9:00 a.m. to 6:00 p.m.