Job Description – Infrastructure Engineer / Senior DevOps Engineer
Job Title: Infrastructure Engineer / Senior DevOps Engineer
About the Role
We are looking for an experienced Infrastructure Engineer / Senior DevOps Engineer (EL3) to design, implement, automate, and maintain secure, scalable, and highly available cloud infrastructure, with a primary focus on Azure, Kubernetes, Terraform, CI/CD, and Observability.
The ideal candidate will have strong hands-on experience with Infrastructure as Code (IaC), AKS, cloud networking, container orchestration, GitHub Actions/Azure DevOps, Datadog, cloud security, FinOps, and DevSecOps.
The role will also involve leveraging enterprise-approved AI tools to automate repetitive tasks, improve engineering productivity, and drive continuous infrastructure improvement.
Key ResponsibilitiesCloud Infrastructure
- Design, implement, and maintain scalable and highly available Azure cloud infrastructure.
- Work with Azure Kubernetes Service (AKS) and containerized workloads.
- Develop and manage cloud networking components including VNets, subnets, NSGs, routing, load balancing, and connectivity.
- Support cloud infrastructure across Azure, AWS, and/or GCP environments.
- Design infrastructure with focus on reliability, scalability, security, and operational efficiency.
Infrastructure as Code
- Develop and maintain infrastructure using Terraform.
- Create reusable Terraform modules and standardized infrastructure patterns.
- Manage infrastructure provisioning, configuration, versioning, and lifecycle.
- Implement Infrastructure as Code best practices across development, staging, and production environments.
- Automate infrastructure changes through CI/CD pipelines.
- Perform Terraform code reviews and ensure consistency, security, and maintainability.
Kubernetes & Containerization
- Deploy, configure, and manage workloads on Kubernetes/AKS.
- Troubleshoot Kubernetes deployments, pods, services, networking, and resource issues.
- Implement container orchestration and scaling strategies.
- Work with Docker/containerized applications.
- Support high availability, rolling deployments, resource management, and cluster optimization.
CI/CD
- Design, implement, and maintain CI/CD pipelines.
- Work with tools such as:
- GitHub Actions
- Azure DevOps
- Jenkins or equivalent
- Automate application and infrastructure deployment.
- Integrate automated testing, security scanning, Terraform validation, and compliance checks into pipelines.
- Implement deployment strategies that improve release reliability and reduce manual intervention.
Observability & Monitoring
- Implement comprehensive cloud and application observability solutions.
- Hands-on experience with Datadog, Prometheus, or similar monitoring platforms.
- Automate the deployment of:
- Monitoring configurations
- Dashboards
- Alerts
- Monitors
- Metrics
- Use Terraform to manage observability infrastructure and configurations.
- Establish monitoring for infrastructure health, application performance, availability, and resource utilization.
- Define meaningful alerting and escalation mechanisms.
- Troubleshoot production incidents using logs, metrics, and traces.
Security & DevSecOps
- Work closely with development and security teams to ensure infrastructure meets organizational security standards.
- Implement DevSecOps practices throughout the software delivery lifecycle.
- Integrate automated security scanning and vulnerability detection into CI/CD pipelines.
- Implement policy-as-code and automated compliance validation.
- Apply cloud security best practices around identity, access control, networking, secrets, and infrastructure configuration.
- Support compliance requirements such as SOC 2, HIPAA, and ISO 27001.
- Identify and remediate infrastructure security vulnerabilities and configuration risks.
Cloud Cost Optimization / FinOps
- Monitor and analyze cloud infrastructure costs.
- Identify opportunities for cost optimization without compromising performance or reliability.
- Implement resource-right-sizing and infrastructure optimization strategies.
- Monitor Kubernetes and cloud resource utilization.
- Apply FinOps principles to improve cloud cost visibility and efficiency.
- Work with engineering and business teams to establish cost-aware infrastructure practices.
Reliability & Disaster Recovery
- Design resilient and highly available cloud infrastructure.
- Implement appropriate backup and disaster recovery strategies.
- Define and support Service Level Objectives (SLOs) and reliability targets.
- Participate in incident management, root-cause analysis, and post-incident reviews.
- Develop automation to improve system resilience and operational recovery.
- Continuously identify and eliminate infrastructure reliability risks.
GitOps
- Implement GitOps practices for infrastructure and application deployments.
- Experience with tools such as:
- Argo CD
- Flux
- Similar GitOps platforms
- Maintain declarative infrastructure and deployment configurations.
- Automate environment synchronization and deployment processes.
AI-Assisted Engineering
- Leverage enterprise-approved AI tools to improve DevOps and infrastructure engineering productivity.
- Use AI tools for:
- Infrastructure code generation
- Terraform assistance
- Troubleshooting
- Documentation
- Automation
- Log analysis
- Test generation
- Operational workflow automation
- Ensure AI-assisted development follows organizational security, privacy, governance, and compliance policies.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent relevant experience.
- 4+ years of experience in DevOps, Cloud Infrastructure, Site Reliability Engineering, or related roles.
- Strong hands-on experience with Terraform and Infrastructure as Code.
- 3+ years of experience working with cloud platforms such as Azure, AWS, and/or GCP.
- Strong experience with Kubernetes, preferably Azure Kubernetes Service (AKS).
- Strong understanding of cloud networking and container orchestration.
- 3+ years of experience developing and managing CI/CD pipelines.
- Hands-on experience with GitHub Actions and/or Azure DevOps.
- Strong understanding of Linux, Git, networking, containers, and cloud infrastructure concepts.
- Strong scripting/automation skills using PowerShell, Bash, Python, or similar languages.
Preferred Qualifications
- Experience implementing Datadog, Prometheus, Grafana, or similar observability platforms.
- Experience automating monitoring, dashboards, alerts, and observability configurations using Terraform.
- Strong knowledge of cloud security and compliance frameworks.
- Experience with SOC 2, HIPAA, ISO 27001, or other regulated environments.
- Experience with FinOps and cloud cost optimization.
- Experience implementing GitOps using Argo CD or Flux.
- Experience with DevSecOps, policy-as-code, vulnerability management, and security scanning.
- Experience designing highly resilient platforms and disaster recovery solutions.
- Experience with Azure services including VNets, AKS, Azure Monitor, Key Vault, Storage, and Azure Policy.
- Experience with AWS/GCP infrastructure is an advantage.
- Relevant certifications such as:
- HashiCorp Terraform Associate
- Microsoft Azure Solutions Architect
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
Key Competencies
- Strong infrastructure architecture and troubleshooting skills.
- Excellent understanding of cloud-native and Kubernetes environments.
- Strong automation and Infrastructure-as-Code mindset.
- Security-first approach to infrastructure engineering.
- Strong analytical and problem-solving capabilities.
- Ability to independently manage complex infrastructure initiatives.
- Strong communication and cross-functional collaboration skills.
- Ability to work effectively with development, security, architecture, and operations teams.
- Strong focus on reliability, scalability, security, observability, and cost efficiency.
Work Location: Remote