Job Description – Senior DevOps / Infrastructure EngineerPosition
Senior DevOps / Infrastructure Engineer
Experience
4+ Years
Job Type
Full-Time
Role Overview
We are looking for an experienced Senior DevOps / Infrastructure Engineer to design, implement, secure, and maintain scalable cloud infrastructure, with a primary focus on Azure, Kubernetes, and Infrastructure as Code (IaC).
The ideal candidate will have strong expertise in Terraform, Kubernetes, CI/CD, cloud networking, observability, security, and cloud cost optimization. You will work closely with development, security, and infrastructure teams to build reliable and secure cloud platforms while leveraging enterprise-approved AI tools to automate workflows and drive continuous improvement.
Key Responsibilities
- Design, implement, and maintain scalable and highly available cloud infrastructure.
- Manage Azure/AWS/GCP environments with a primary focus on Azure.
- Build and maintain infrastructure using Terraform and Infrastructure as Code (IaC) best practices.
- Design, configure, and manage Kubernetes environments, particularly Azure Kubernetes Service (AKS).
- Develop and maintain CI/CD pipelines using GitHub Actions, Azure DevOps, or similar tools.
- Implement comprehensive observability and monitoring solutions using Datadog, Prometheus, or equivalent platforms.
- Automate the deployment of Datadog monitors, dashboards, alerts, and observability components using Terraform.
- Monitor infrastructure performance, availability, reliability, and capacity.
- Collaborate with security teams to ensure infrastructure and applications comply with organizational security standards.
- Implement DevSecOps practices, including automated security scanning, vulnerability management, policy-as-code, and compliance validation.
- Monitor cloud infrastructure costs and identify opportunities for cost optimization.
- Apply FinOps principles to improve cloud resource utilization and reduce unnecessary spending.
- Design resilient infrastructure with appropriate disaster recovery, backup, reliability, and business continuity strategies.
- Define and maintain Service-Level Objectives (SLOs) and reliability standards.
- Implement GitOps practices using tools such as Argo CD, Flux, or similar technologies.
- Leverage enterprise-approved AI tools to streamline engineering workflows, automate repetitive tasks, and improve operational efficiency.
- Troubleshoot complex infrastructure, deployment, networking, and production issues.
- Work collaboratively with technical and non-technical stakeholders to deliver reliable cloud solutions.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent relevant experience.
- 4+ years of experience in DevOps or Infrastructure Engineering.
- Strong hands-on experience with Terraform and Infrastructure as Code.
- 3+ years of experience working with cloud platforms such as Azure, AWS, or GCP.
- Strong experience with Kubernetes, preferably Azure Kubernetes Service (AKS).
- Good understanding of cloud networking, virtual networks, containers, and orchestration.
- 3+ years of experience developing and managing CI/CD pipelines.
- Hands-on experience with tools such as GitHub Actions and Azure DevOps.
- Strong troubleshooting, automation, and problem-solving skills.
Preferred Qualifications
- Experience implementing and managing Datadog, Prometheus, or similar observability platforms.
- Experience automating monitoring, dashboards, alerts, and observability infrastructure.
- Strong understanding of cloud security best practices and compliance frameworks, including SOC 2, HIPAA, or ISO 27001.
- Knowledge of cloud security controls and compliance requirements.
- Experience with cloud cost optimization and FinOps principles.
- Experience with GitOps using Argo CD, Flux, or similar tools.
- Strong knowledge of DevSecOps practices, including:
- Automated security scanning
- Policy-as-code
- Vulnerability management
- Compliance validation
- Experience designing highly resilient cloud platforms.
- Knowledge of disaster recovery, backup strategies, high availability, and SLOs.
- Experience working in regulated industries such as healthcare, financial services, or other compliance-driven environments.
- Relevant certifications such as:
- HashiCorp Terraform Associate
- Microsoft Azure Solutions Architect
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
Technical Skills
Cloud: Azure, AWS, GCP Infrastructure as Code: Terraform Containers & Orchestration: Kubernetes, AKS, Docker CI/CD: GitHub Actions, Azure DevOps Observability: Datadog, Prometheus GitOps: Argo CD, Flux DevSecOps: Security Scanning, Policy-as-Code, Vulnerability Management Cloud Security: IAM, Security Controls, Compliance Cloud Optimization: FinOps, Cost Monitoring & Optimization Reliability: Disaster Recovery, Backup, High Availability, SLOs
Key Competencies
- Strong analytical and problem-solving abilities
- Excellent communication and collaboration skills
- Strong automation mindset
- Ability to work across development, security, and infrastructure teams
- Strong focus on reliability, security, and operational excellence
- Ability to identify and implement continuous improvement opportunities
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)