
1
applicant
applicant
Senior Data Engineer - Azure Databricks / PySpark
EFutures
Colombo •
Full-time
Job Description
Join our team and work on exciting data and MDM initiatives building scalable data pipelines and data solutions for enterprise systems.
Key Responsibilities
- Develop ingestion pipelines for REST APIs, database/CDC and file-based sources.
- Build PySpark/Spark transformations and Delta Lake processing on Azure Databricks.
- Implement Bronze/raw, standardised/conformed and governed data patterns as defined by architecture.
- Develop reusable standardisation, validation and DQ processing before MDM.
- Implement incremental loading, schema validation/evolution and restart/recovery patterns.
- Create source-to-canonical transformations while retaining lineage/provenance.
- Optimise Spark jobs for performance, cost and operational reliability.
- Write automated unit/data-quality tests and participate in peer review.
- Integrate Databricks outputs/inputs with Profisee and downstream Azure services.
- Support production monitoring, troubleshooting and runbook creation.
Must-Have Experience & Skills
- Overall 7+ years of experience and 5+ years data engineering with strong production PySpark/Spark experience.
- Strong Azure Databricks, Delta Lake, SQL and ADLS Gen2.
- REST/API ingestion and database integration experience.
- Incremental/CDC processing and schema-evolution experience.
- Strong data-quality/testing practices and Git/CI/CD.
- Experience handling large, messy enterprise datasets.
Preferred/Nice-to-Have
- Databricks certification.
- MDM/entity-resolution project exposure.
- Azure Data Factory or equivalent orchestration.
- Exposure to ML-based entity resolution / fuzzy matching (PySpark MLlib).
- Purview/Unity Catalog governance exposure.
Employment Details
- Employment Type: Part Time (2 Years Contract)
- Work Arrangement: Hybrid (3 to 4 days onsite)
- Working Days: Monday to Friday
- Working Hours: 9:00 a.m. to 6:00 p.m.