Job Description – Cloud Platform Engineer
Job Title: Cloud Platform Engineer
About the Role
We are seeking a skilled Cloud Platform Engineer (EL2) to join our team for a Cloud Data Modernization initiative. The ideal candidate will have strong hands-on experience with Databricks as the primary data platform, supported by Snowflake and Azure/AWS cloud infrastructure.
The role involves modernizing and migrating on-premises ETL workloads to cloud-based data platforms, developing scalable data pipelines, implementing CI/CD practices, and ensuring strong data quality, observability, security, and performance.
The successful candidate should be comfortable working with Databricks, Apache Spark, PySpark, Azure Data Factory, ADLS Gen2, Snowflake, Python, SQL, and GitHub Actions.
Key Responsibilities
- Contribute to the migration and modernization of on-premises ETL workloads to cloud platforms.
- Develop and maintain scalable data engineering solutions using Databricks as the primary platform.
- Build data pipelines using Databricks, Apache Spark, PySpark, and Azure Data Factory (ADF).
- Work with Snowflake as a secondary data platform for data engineering and analytics workloads.
- Develop and support integrations using ADF, Self-hosted Integration Runtime (SHIR), Logic Apps, ADLS Gen2, and Blob Storage.
- Review existing on-premises ETL processes and identify opportunities for cloud modernization and optimization.
- Implement Bronze/Silver/Gold (Medallion) architecture and Lakehouse patterns where applicable.
- Work with Delta Lake for scalable and reliable data processing.
- Develop data transformations and ETL/ELT processes using Python, PySpark, and SQL.
- Implement automated data validation, reconciliation, data quality, monitoring, and observability processes.
- Optimize Databricks and Snowflake workloads for performance, scalability, reliability, and cost efficiency.
- Implement and maintain CI/CD pipelines using GitHub Actions.
- Follow Git branching, pull-request, code-review, and engineering best practices.
- Collaborate with data engineers, cloud engineers, architects, QA teams, business analysts, and other stakeholders.
- Support production data pipelines and troubleshoot data, performance, integration, and deployment issues.
- Ensure data solutions follow applicable security, governance, compliance, and data-quality standards.
- Leverage AI-assisted development tools to improve engineering productivity, code quality, testing, and documentation.
Required Technical SkillsDatabricks – Primary Skill
- Strong hands-on experience with Databricks.
- Experience developing data pipelines using PySpark and Apache Spark.
- Good understanding of Delta Lake and Lakehouse architecture.
- Experience with Databricks notebooks, jobs/workflows, clusters, and optimization.
- Understanding of Medallion Architecture – Bronze, Silver, and Gold layers.
- Experience optimizing Spark jobs, transformations, partitions, joins, and workloads.
Azure Cloud
Hands-on experience with Azure data services such as:
- Azure Data Factory (ADF)
- Self-hosted Integration Runtime (SHIR)
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Blob Storage
- Azure Logic Apps
Good understanding of cloud-based data integration and storage architectures.
Snowflake – Secondary Skill
- Hands-on experience with Snowflake for data engineering and analytics workloads.
- Experience developing SQL-based transformations and ETL/ELT processes.
- Understanding of Snowflake data loading, warehouse concepts, and performance optimization.
- Experience supporting Snowflake migration or modernization initiatives is desirable.
Programming & Data Engineering
- Strong experience with Python and/or PySpark.
- Strong SQL development skills.
- Experience writing complex queries, joins, stored procedures, views, CTEs, and data transformations.
- Experience developing enterprise ETL/ELT pipelines.
- Understanding of structured and semi-structured data processing.
DevOps & CI/CD
- Hands-on experience with Git and GitHub.
- Experience with GitHub branching strategies and Pull Requests.
- Experience developing and maintaining GitHub Actions CI/CD pipelines.
- Understanding of automated testing, deployment, and release processes.
- Experience participating in code reviews and following engineering best practices.
Data Quality & Observability
- Experience implementing data validation and reconciliation frameworks.
- Knowledge of data quality dimensions such as accuracy, completeness, consistency, and timeliness.
- Experience with pipeline monitoring and data observability.
- Ability to identify and troubleshoot data quality issues.
Cloud Platform Experience
- Strong experience with Azure is preferred.
- Experience with AWS is an advantage.
- Understanding of cloud storage, compute, networking, security, and identity concepts.
- Ability to work across cloud environments and adapt to new cloud technologies.
AI-Assisted Engineering
- Experience using AI-assisted development tools such as:
- GitHub Copilot
- Databricks Assistant
- ChatGPT
- Claude
- Similar AI coding assistants
- Ability to use AI tools for code generation, debugging, refactoring, test development, documentation, and engineering productivity while following enterprise security and governance standards.
Preferred Qualifications
- Experience with Delta Lake and Lakehouse architectures.
- Experience implementing Bronze/Silver/Gold architecture.
- Experience with Airflow or other orchestration frameworks.
- Experience with data modeling and database design.
- Knowledge of data governance and data quality best practices.
- Experience with Databricks and Snowflake performance optimization.
- Experience with healthcare payer data such as:
- Claims
- Membership
- Enrollment
- Provider
- Clinical
- Financial data
- Experience leveraging AI/ML for data engineering automation.
- Experience with other cloud platforms and data solutions.
- Azure or Databricks certification is a plus.
Key Responsibilities by PriorityPrioritySkill / AreaPrimary DatabricksPrimary PySpark / Apache SparkPrimary Cloud Data EngineeringPrimary Azure Data PlatformSecondary SnowflakeSecondary SQL / PythonSecondary ETL/ELT & Data MigrationRequired GitHub / GitHub Actions / CI/CDRequired Data Quality & ValidationGood to Have Delta Lake / LakehouseGood to Have Medallion ArchitectureGood to Have AirflowGood to Have Healthcare Payer DataGood to Have AI-Assisted DevelopmentKey Competencies
- Strong analytical and problem-solving skills.
- Ability to work independently and manage ambiguity.
- Strong understanding of cloud data engineering concepts.
- Ability to troubleshoot complex data and pipeline issues.
- Strong focus on data quality, security, scalability, and performance.
- Excellent communication and collaboration skills.
- Ability to work effectively within Agile delivery methodologies.
- Ability to learn and adopt new technologies quickly.
- Ability to work under tight deadlines and proactively escalate critical issues.
Work Location: Remote