← All jobs
Apple

Senior ML Infrastructure SRE

Austin, TX, USOn-siteSeniorvia jobspy_indeed
kubernetescloudautomationsrepythonstoragelinuxobservability

Observed pay for machine learning engineer jobs in the united states

Among 19 openings that publish compensation, the median stated annual range is $132K–$200K USD. This uses observed listing data across 405 live jobs, never an estimated salary.

Don't apply into the void.

Most applications for this Apple role vanish into an ATS. With jobfinder-ai, your agent finds the actual hiring manager or founder behind this opening and sends a tailored email from your own inbox — so a real person reads your pitch and replies. We then follow up until you land on the calendar.

Reach the decision-maker — $5

View original posting →

At Apple, we don’t just build products \- we create transformative experiences that have reshaped entire industries. Our innovation is driven by the diversity of our people and their ideas, inspiring everything we do. Imagine the impact you could make. Join Apple and help us leave the world better than we found it. The ML Infrastructure team is responsible for managing Apple’s largest ML compute platform, multi\-cloud storage abstraction and caching platform, which supports critical machine learning training workloads that power user\-facing features across the Apple ecosystem. Operating across both first\-party and third\-party cloud environments brings complex and unique challenges. As a Site Reliability Engineer (SRE) on the ML Infrastructure team, you’ll be expected to address these challenges through a strong foundation in cloud object storage, data analysis, automation, collaboration, and advanced expertise in Kubernetes. Our team oversees the full infrastructure stack \- from low\-level nodes to the complete network architecture \- ensuring our platform remains highly available, resilient, and efficient at scale.

**Description**

We are seeking an experienced Software and Systems Engineer to join our dynamic team. This role demands a proactive mindset, technical excellence, and a collaborative spirit. The ideal candidate will demonstrate: • Strong critical thinking and a high degree of individual accountability • Effective communication and collaboration skills • A genuine passion for Infrastructure as a Service (IaaS) • A commitment to automation and operational efficiency • Ownership of projects from design through delivery • A solutions\-oriented approach, coupled with the ability to gain alignment on technical direction • Consistent and timely execution of design implementations aligned with project objectives • The ability to provide constructive technical feedback, fostering team\-wide growth and continuous improvement

","responsibilities":"Participates in a rotating on\-call schedule, including occasional weekend

coverage when necessary

Currently headquartered in Cupertino, with active expansion in Bangalore to

support global operations across time zones

Leverages a diverse stack including open\-source tools, commercial solutions,

and internally developed systems

Encourages open dialogue, values strong ideas, and recognizes impactful results

**Preferred Qualifications**

Proven drive to automate manual operations and enhance processes through

continuous iteration

Strong understanding of best practices for deploying large\-scale, distributed

applications

Hands\-on experience managing diverse system environments using

configuration management tools or software delivery platforms such as

Spinnaker, Helm, or Flux

Demonstrated expertise in deploying, supporting, and monitoring both new

and existing services, platforms, and application stacks

Solid familiarity with container orchestration and management using

Kubernetes

**Minimum Qualifications**

5\+ years experience in building, operating and scaling a large application in a

private, public or hybrid cloud environment

Deep expertise in Kubernetes, with hands\-on experience using platforms such

as Google Kubernetes Engine (GKE) or Amazon Elastic Kubernetes Service

(EKS)

Proficient in designing, developing, and releasing code in languages such as

Python, Go, or Rust

Practical experience with object storage technologies, including Amazon S3

or Google Cloud Storage (GCS)

Strong background in designing and troubleshooting complex networking

issues in both public and private cloud infrastructures

Solid understanding of Linux internals, standard networking protocols, and

distributed systems architecture

Set this role as a target and your agent does the sourcing, finds the verified email, writes the pitch, and follows up — on autopilot.

Start your hunt