Browse more devops engineer jobs
Manager, Core Infrastructure Engineering
Salary context from live listings
Observed pay for devops engineer jobs in the united states
Among 25 openings that publish compensation, the median stated annual range is $100K–$150K USD. This uses observed listing data across 276 live jobs, never an estimated salary.
Don't apply into the void.
Most applications for this Oracle role vanish into an ATS. With jobfinder-ai, your agent finds the actual hiring manager or founder behind this opening and sends a tailored email from your own inbox — so a real person reads your pitch and replies. We then follow up until you land on the calendar.
Reach the decision-maker — $5About the role
For a single team delivering components of distributed systems. Translates goals into a 1–2 quarter execution plan, sets coding, testing, and scalability practices, and provides hands\-on oversight of performance tuning and load/perf testing. Guides the team in building fault\-tolerant, in\-service\-upgradable components (redundancy, replication, failover) and in applying resiliency patterns (retries, circuit breakers, timeouts). Ensures robust observability (tests, alarms, dashboards, telemetry) and operational readiness via reviewed runbooks and standard procedures. Manages delivery of scoped features and correctness testing (including fault\-injection/brownouts), and directs implementation of data replication/synchronization to maintain integrity and availability. Leads team incident response and root\-cause efforts, enforces no\-customer\-downtime practices, and drives use of automation/IaC for troubleshooting. Oversees team security implementation (encryption, access controls), tracks remediation plans, verifies compliance documentation, and coaches adherence to change\-management plans for safe patching, updates, and rollbacks.
**Key Responsibilities**
**System Design \& Architecture – System Scalability:**
* Provides oversight to the team on the development of components of distributed systems, including the use of distributed state management tools. * Monitors and enables optimization of code and/or systems for large\-scale data processing. * Guides team to implement scalability requirements for assigned components. * Coaches team to leverage data plane platforms to effectively handle large\-scale data retrieval, storage, and processing. * Ensures team accurately implements and executes performance and load testing.
**System Design \& Architecture – System Reliability Design:**
* Provides guidance to the team in building fault\-tolerant systems capable of withstanding in\-service updates by leading the implementation of redundancy, replication, and automatic failover mechanisms. * Guides the design of components to effectively handle service disruptions. * Coaches team on various approaches to handle network unreliability, including retry mechanisms, circuit breakers, and timeouts.
**System Design \& Architecture – System Reliability Performance:**
* Guides team to implement testing and alarming configurations for detecting and addressing issues/failures. * Ensures team effectively supports recovery efforts by reviewing and aligning on runbooks and operational procedures. * Coaches team on building and customizing dashboards, telemetry systems, and alerting mechanisms to monitor component health.
**System Design \& Architecture – Correctness / Availability:**
* Manages the implementation and design of functional requirements and testing for features within an existing system. * Coaches team on implementing test scenarios (e.g., fault\-injection, brown\-out) to evaluate system correctness. * Guides the implementation of data replication and synchronization techniques within the team to maintain data integrity and availability.
**Operational Troubleshooting \& Incident Management:**
* Leads team efforts in diagnosing, debugging, and resolving issues in system components to support ongoing operation. * Ensures teams are following protocols for preventing interruptions, ensuring no maintenance windows are required for customers and users when resolving issues. * Manages team in effectively implementing automation scripts and tooling when troubleshooting operational issues. * Creates schedules and manages operational support rotations.
**Compliance \& Security:**
* Manages team implementation of robust security measures to protect data and applications in multi\-tenant environments, overseeing encryption techniques and access controls. * Manages execution of remediation plans to address identified security gaps, ensuring continuous improvement of security measures. * Reviews documentation and ensures cloud infrastructure compliance with industry standards and regulations.
**Automation \& Change Management:**
* Leads the maintenance of automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure. * Coaches team on change management plans for patching, updating, and rolling back applications.
**Core Responsibilities**
**Planning \& Execution:**
* Creates and owns the execution plan for the team’s work and multiple projects or initiatives, monitoring timelines and budgets (when applicable) to ensure projects are completed on time and in adherence with requirements. * Delegates work across the team and helps them prioritize their work. * Identifies resource needs and adapts plans based on changing priorities and business needs.
**Collaboration \& Partnership:**
* Strengthens collaborative partnerships across teams to align on expectations and shared objectives. * Guides team members to build relationships with business leaders, stakeholders, and/or customers to ensure effective collaboration. * Practices active listening and asks insightful questions to promote an inclusive culture.
**Problem Solving:**
* Leads team to identify and address moderately complex operational and/or technical issues in accordance with standard practices, providing guidance as appropriate. * Directs team to analyze data and/or information from multiple sources to troubleshoot moderately complex errors.
**Continuous Learning:**
* Actively seeks learning opportunities for self and team to enhance knowledge and skills in key areas and remain current with industry advancements. * Uses feedback and training to elevate personal and team skills, modeling a commitment to learning. * Identifies skill gaps within the team and supports team members in learning opportunities by providing resources and fostering an environment that encourages knowledge\-sharing.
**Continuous Improvement:**
* Identifies and recommends improvements and encourages the team to share ideas to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. * Reviews and provides feedback on team members’ ideas and seeks input from others on alternative approaches and methods for improving work.
**Performance and Development:**
* Coaches and provides direction to team members in alignment with performance management processes, guidelines, and expectations. * Facilitates goal\-setting discussions with team members, ensuring individual goals are aligned with broader team goals. * Identifies development opportunities to help team members improve performance. * Contributes to the talent acquisition pipeline by leading candidate interviews, assessing promotion eligibility, and managing talent resources.
Ready to reach the decision-maker?
Set this role as a target and your agent does the sourcing, finds the verified email, writes the pitch, and follows up — on autopilot.
Start your hunt